Kembali ke Berita
produkAI Understanding taklimat

NVIDIA says Vera Rubin NVL72 delivers up to 30x higher agentic-AI throughput per megawatt

NVIDIA reports that its Vera Rubin NVL72 systems produced up to 30x more agentic-AI inference throughput per megawatt and up to 35x lower token costs than GB300 NVL72 systems in company measurements. The results used recorded coding-agent trajectories and remain pending SemiAnalysis review.

6 min readRead the primary source
Primary-source image accompanying NVIDIA says Vera Rubin NVL72 delivers up to 30x higher agentic-AI throughput per megawatt
Dokumen sumber utamaSumber direkodkan
Penerbit
blogs.nvidia.com
Pautan sumber
blogs.nvidia.comhttps://blogs.nvidia.com/blog/vera-rubin-nvl72-efficiency-ai-agents/
Jenis sumber
Dokumen utama — pengumuman rasmi, kertas, pemfailan atau halaman pihak pertama yang kami baca secara langsung.
Juga dipetik

Cerita terakhir disemak

KonteksFahami perkara ini dalam masa 60 saat

Mulakan di sini

Istilah utama

Kuantisasi
Menukar pemberat model kepada format ketepatan yang lebih rendah seperti 8-bit atau 4-bit.
Transformer
Seni bina saraf yang menggunakan perhatian kepada perhubungan model merentas jujukan secara selari.
Penanda aras
Ujian piawai atau set data yang digunakan untuk mengukur dan membandingkan prestasi model.
Uji diri andaKuiz Agen AI

Apa yang berubah sejak penerbitan

  1. Pertama kali diterbitkan
  2. This source materially advances the same Vera Rubin and Blackwell agentic-inference efficiency event already covered by the canonical update. It adds NVIDIA’s detailed claim of up to 30x higher throughput per megawatt and 35x lower token cost for Vera Rubin NVL72 versus GB300 NVL72, identifies the SemiAnalysis AgentX workload and its preserved coding-agent trajectories, describes the claimed hardware-software optimizations, and explicitly notes that the results are early and pending SemiAnalysis review and exclude Vera CPU performance for tool calling.
  3. This is a continuing update to NVIDIA’s reported agentic-inference efficiency results. The new post adds methodological detail about SemiAnalysis AgentX, including replayed Claude Code sessions, context reuse, tool-call timing, concurrency and user-experience metrics. It reiterates the headline claim that Vera Rubin NVL72 delivered up to 30x higher throughput per megawatt than GB300 NVL72, while explicitly noting that NVIDIA measured the result and that SemiAnalysis review is still pending.
  4. This NVIDIA blog is a same-day primary-source elaboration of the existing Vera Rubin efficiency report. It adds details about the AgentX coding trajectories, preserved context growth, tool calls and sub-agent spawning, while also stating that the results are early, measured by NVIDIA, pending SemiAnalysis review and incomplete because Vera CPU tool-calling performance is not included.

Apa yang berlaku

NVIDIA reports new performance measurements for Vera Rubin NVL72 systems running agentic-AI workloads, claiming substantially higher throughput per megawatt and lower token costs than its GB300 NVL72 platform.

NVIDIA said on August 24 that Vera Rubin NVL72 systems delivered up to 30 times higher agentic-AI inference throughput per megawatt than NVIDIA GB300 NVL72 systems. The company also claimed up to 35 times lower cost per million tokens. These are NVIDIA’s figures, not independently established results, and the source presents them as early performance data rather than a general guarantee for every workload. The comparison is specifically between NVIDIA’s Vera Rubin and GB300 NVL72 systems.

The measurements used SemiAnalysis AgentX, which NVIDIA describes as a workload built from recorded real-world agentic coding sessions. According to the source, the preserves context growth, tool calls and the spawning of sub-agents. NVIDIA says the workload was intended to measure an entire agent workflow rather than a single request. The source lists Kimi K3, MiniMax M3, GLM5.3, Qwen3.5 and DeepSeek V4 Pro among the models used in the broader AgentX results, and gives DeepSeek V4 Pro as an example where Vera Rubin reached as much as 30 times the throughput per megawatt of GB300 NVL72.

NVIDIA frames agentic workloads as unusually demanding because an agent may repeatedly retrieve information, call tools, invoke other agents and carry accumulated context into later steps. The source contrasts these sessions with chat or document-summarization requests that typically involve input and output sequences of 1,000 to 8,000 tokens. It says agentic sessions can reach hundreds of thousands of input tokens, with substantial variation in input and output lengths. NVIDIA also cites OpenRouter data saying agentic workloads consume 15 times more tokens than a simple chat request, though the source does not provide the underlying data or methodology.

The company attributes the reported gains to a rack-scale platform and software stack that combine disaggregated prefill and decode serving, rate matching, expert parallelism, distributed key-value caching, cache-aware routing, fused CUDA kernels, Tensor Cores, Engine, NVFP4 and high-bandwidth NVLink interconnects. NVIDIA says the full Vera Rubin platform includes seven chip types, including the Vera CPU, Groq 3 LPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX and ConnectX-9 SuperNIC. The source says Vera Rubin is in full production and scaling across its ecosystem, but it does not specify customer deployments, volumes, prices or general availability terms.

Butiran sumber: blogs.nvidia.com

Mengapa ia penting

If independently confirmed across varied workloads, the claimed efficiency gains could affect the energy, operating-cost and infrastructure requirements of large-scale AI agents. The evidence is currently limited to NVIDIA’s measurements and a workload review that is still pending.

The central significance of the announcement is energy efficiency. NVIDIA argues that throughput per megawatt can determine how much agentic work an AI factory performs within a fixed power budget. If the reported comparison survives independent review, operators with constrained electricity capacity could potentially run more agent sessions without increasing their power allocation. That could matter as agents perform longer sequences of retrieval, reasoning and tool use than ordinary chat systems. The source does not establish how the result would translate into savings at a particular facility.

The claimed token-cost reduction matters because agent systems may generate and process much more context than a single-turn application. NVIDIA says cost per million tokens affects the profit margin of AI-factory operators. Lower inference cost could make some long-running coding, research or customer-service workflows easier to operate continuously. That conclusion remains conditional: the source gives relative results, not an absolute cost per token, a full system price, deployment cost, electricity price or total cost of ownership.

The measurement approach is notable because it attempts to represent agent behavior over multiple steps. A based only on one prompt and one response could miss the overhead created by repeated tool calls, growing context and coordination among sub-agents. NVIDIA says AgentX retains those features from recorded coding sessions. That makes the result more relevant to the type of workloads the company is targeting, but it also means the outcome may depend heavily on the selected trajectories, models, context lengths, serving configuration and software optimizations.

There are material limits to the evidence. The results were measured by NVIDIA, the page says they are pending SemiAnalysis review, and the headline figures use “up to,” which describes a maximum rather than an average. The source does not report confidence intervals, the number of trajectories, absolute throughput, latency, quality comparisons, failure rates or results from independent operators. It also says the measurements do not yet reflect Vera CPU performance for tool calling. Readers therefore cannot conclude from this source alone that every agent workload will be 30 times more efficient or that output quality will remain equivalent in all conditions.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Interactive Concept Check+10 Points
AI Agents Quiz

What is the most accurate way to describe what AI Agents can do today?

Apa yang perlu ditonton seterusnya

The key next checks are SemiAnalysis’s review, independent testing across models and agent tasks, results that include Vera CPU tool-calling performance, and evidence that the claimed gains hold beyond peak configurations and selected trajectories.

The first point to watch is the SemiAnalysis review. It could clarify how AgentX was constructed, how many coding trajectories were measured, which settings were used, how throughput and power were calculated, and whether the comparison can be reproduced. The review may also show whether the headline result reflects a typical performance range or a best-case point. Until those details are available, NVIDIA’s figures should be treated as company-reported measurements.

Independent results across different agent tasks will be important. Coding trajectories may not represent research, customer service, data analysis or other workflows with different tool patterns and context-growth behavior. Future testing should show whether the relative advantage persists across models, request lengths, sub-agent counts, cache behavior and quality targets. The source names several models, but it does not provide a complete workload-by-workload table or establish that the claimed maximum applies broadly.

The source specifically says that the results do not yet include Vera CPU performance for tool calling. That component could be important for agents whose efficiency depends on orchestration, data access and other sequential operations outside GPU inference. Further reporting should separate GPU inference results from end-to-end agent performance and show how the Vera CPU and the other platform components affect latency, utilization and power consumption during complete workflows.

NVIDIA also claims that DSX MaxLPS can provision up to 40% more GPUs within the same megawatt budget and that NVLink 6 delivers higher packet rates and lower latency than off-the-shelf Ethernet alternatives. Those claims warrant measurements that include facility-level power behavior, cooling, networking and utilization rather than only accelerator throughput. The source also says NVFP4 increases throughput without sacrificing output quality, but it supplies no quality tests here. Independent evaluation of accuracy, reliability, latency, availability and total operating cost will determine how much of the claimed efficiency translates into practical public or commercial impact.

Panduan & kuiz berkaitan

Ejen AIModel AI DiterangkanLatihan AIMasa Depan AIUji apa yang anda tahu — cuba kuiz AI percumaCari istilah AI dalam glosari kami

Kemas kini dan pembetulan

Kisah kanonik ini dikemas kini apabila peristiwa yang sedang berkembang berubah secara material. URL dan tarikh penerbitan asalnya tidak pernah berubah.

  • This NVIDIA blog is a same-day primary-source elaboration of the existing Vera Rubin efficiency report. It adds details about the AgentX coding trajectories, preserved context growth, tool calls and sub-agent spawning, while also stating that the results are early, measured by NVIDIA, pending SemiAnalysis review and incomplete because Vera CPU tool-calling performance is not included.
  • This is a continuing update to NVIDIA’s reported agentic-inference efficiency results. The new post adds methodological detail about SemiAnalysis AgentX, including replayed Claude Code sessions, context reuse, tool-call timing, concurrency and user-experience metrics. It reiterates the headline claim that Vera Rubin NVL72 delivered up to 30x higher throughput per megawatt than GB300 NVL72, while explicitly noting that NVIDIA measured the result and that SemiAnalysis review is still pending.
  • This source materially advances the same Vera Rubin and Blackwell agentic-inference efficiency event already covered by the canonical update. It adds NVIDIA’s detailed claim of up to 30x higher throughput per megawatt and 35x lower token cost for Vera Rubin NVL72 versus GB300 NVL72, identifies the SemiAnalysis AgentX workload and its preserved coding-agent trajectories, describes the claimed hardware-software optimizations, and explicitly notes that the results are early and pending SemiAnalysis review and exclude Vera CPU performance for tool calling.
Lihat log pembetulan awam
Adakah ini berguna?