Ku laabo Warka
AlaabtaAI Understanding warbixin kooban

NVIDIA says Vera Rubin NVL72 delivers up to 30x higher agentic-AI throughput per megawatt

NVIDIA reports that its Vera Rubin NVL72 systems produced up to 30x more agentic-AI inference throughput per megawatt and up to 35x lower token costs than GB300 NVL72 systems in company measurements. The results used recorded coding-agent trajectories and remain pending SemiAnalysis review.

6 min readRead the primary source
Primary-source image accompanying NVIDIA says Vera Rubin NVL72 delivers up to 30x higher agentic-AI throughput per megawatt
Dukumeentiga isha aasaasiga ahIsha la duubay
Daabacaha
blogs.nvidia.com
Xidhiidhka isha
blogs.nvidia.comhttps://blogs.nvidia.com/blog/vera-rubin-nvl72-efficiency-ai-agents/
Nooca isha
Dukumeentiga aasaasiga ah - ogeysiis rasmi ah, warqad, xereyn, ama bogga xisbiga koowaad waxaan si toos ah u akhrinay.
Sidoo kale la soo xigtay

Sheekada ayaa dib loo eegay

Dulucda sheekadaKu fahan tan 60 ilbiriqsi gudahood

Halkan ka bilow

Qodobbada muhiimka ah

Tirada
Miisaanka moodeelka oo loo beddelo qaababka saxda ah ee hoose sida 8-bit ama 4-bit.
Transformer
Nashqada neerfaha ee adeegsata fiiro gaar ah u qaabaynta xidhiidhada isku xigxiga ee is barbar socda.
Benchmark
Tijaabo la habeeyey ama kayd xogeed oo loo isticmaalo in lagu cabbiro laguna barbar dhigo waxqabadka moodeelka.
Is tijaabiKediska Wakiilada AI

Maxaa isbedelay tan iyo markii la daabacay

  1. Marka hore la daabacay
  2. This source materially advances the same Vera Rubin and Blackwell agentic-inference efficiency event already covered by the canonical update. It adds NVIDIA’s detailed claim of up to 30x higher throughput per megawatt and 35x lower token cost for Vera Rubin NVL72 versus GB300 NVL72, identifies the SemiAnalysis AgentX workload and its preserved coding-agent trajectories, describes the claimed hardware-software optimizations, and explicitly notes that the results are early and pending SemiAnalysis review and exclude Vera CPU performance for tool calling.
  3. This is a continuing update to NVIDIA’s reported agentic-inference efficiency results. The new post adds methodological detail about SemiAnalysis AgentX, including replayed Claude Code sessions, context reuse, tool-call timing, concurrency and user-experience metrics. It reiterates the headline claim that Vera Rubin NVL72 delivered up to 30x higher throughput per megawatt than GB300 NVL72, while explicitly noting that NVIDIA measured the result and that SemiAnalysis review is still pending.
  4. This NVIDIA blog is a same-day primary-source elaboration of the existing Vera Rubin efficiency report. It adds details about the AgentX coding trajectories, preserved context growth, tool calls and sub-agent spawning, while also stating that the results are early, measured by NVIDIA, pending SemiAnalysis review and incomplete because Vera CPU tool-calling performance is not included.

Maxaa dhacay

NVIDIA reports new performance measurements for Vera Rubin NVL72 systems running agentic-AI workloads, claiming substantially higher throughput per megawatt and lower token costs than its GB300 NVL72 platform.

NVIDIA said on August 24 that Vera Rubin NVL72 systems delivered up to 30 times higher agentic-AI inference throughput per megawatt than NVIDIA GB300 NVL72 systems. The company also claimed up to 35 times lower cost per million tokens. These are NVIDIA’s figures, not independently established results, and the source presents them as early performance data rather than a general guarantee for every workload. The comparison is specifically between NVIDIA’s Vera Rubin and GB300 NVL72 systems.

The measurements used SemiAnalysis AgentX, which NVIDIA describes as a workload built from recorded real-world agentic coding sessions. According to the source, the preserves context growth, tool calls and the spawning of sub-agents. NVIDIA says the workload was intended to measure an entire agent workflow rather than a single request. The source lists Kimi K3, MiniMax M3, GLM5.3, Qwen3.5 and DeepSeek V4 Pro among the models used in the broader AgentX results, and gives DeepSeek V4 Pro as an example where Vera Rubin reached as much as 30 times the throughput per megawatt of GB300 NVL72.

NVIDIA frames agentic workloads as unusually demanding because an agent may repeatedly retrieve information, call tools, invoke other agents and carry accumulated context into later steps. The source contrasts these sessions with chat or document-summarization requests that typically involve input and output sequences of 1,000 to 8,000 tokens. It says agentic sessions can reach hundreds of thousands of input tokens, with substantial variation in input and output lengths. NVIDIA also cites OpenRouter data saying agentic workloads consume 15 times more tokens than a simple chat request, though the source does not provide the underlying data or methodology.

The company attributes the reported gains to a rack-scale platform and software stack that combine disaggregated prefill and decode serving, rate matching, expert parallelism, distributed key-value caching, cache-aware routing, fused CUDA kernels, Tensor Cores, Engine, NVFP4 and high-bandwidth NVLink interconnects. NVIDIA says the full Vera Rubin platform includes seven chip types, including the Vera CPU, Groq 3 LPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX and ConnectX-9 SuperNIC. The source says Vera Rubin is in full production and scaling across its ecosystem, but it does not specify customer deployments, volumes, prices or general availability terms.

Faahfaahinta isha: blogs.nvidia.com

Maxay muhiim u tahay

If independently confirmed across varied workloads, the claimed efficiency gains could affect the energy, operating-cost and infrastructure requirements of large-scale AI agents. The evidence is currently limited to NVIDIA’s measurements and a workload review that is still pending.

The central significance of the announcement is energy efficiency. NVIDIA argues that throughput per megawatt can determine how much agentic work an AI factory performs within a fixed power budget. If the reported comparison survives independent review, operators with constrained electricity capacity could potentially run more agent sessions without increasing their power allocation. That could matter as agents perform longer sequences of retrieval, reasoning and tool use than ordinary chat systems. The source does not establish how the result would translate into savings at a particular facility.

The claimed token-cost reduction matters because agent systems may generate and process much more context than a single-turn application. NVIDIA says cost per million tokens affects the profit margin of AI-factory operators. Lower inference cost could make some long-running coding, research or customer-service workflows easier to operate continuously. That conclusion remains conditional: the source gives relative results, not an absolute cost per token, a full system price, deployment cost, electricity price or total cost of ownership.

The measurement approach is notable because it attempts to represent agent behavior over multiple steps. A based only on one prompt and one response could miss the overhead created by repeated tool calls, growing context and coordination among sub-agents. NVIDIA says AgentX retains those features from recorded coding sessions. That makes the result more relevant to the type of workloads the company is targeting, but it also means the outcome may depend heavily on the selected trajectories, models, context lengths, serving configuration and software optimizations.

There are material limits to the evidence. The results were measured by NVIDIA, the page says they are pending SemiAnalysis review, and the headline figures use “up to,” which describes a maximum rather than an average. The source does not report confidence intervals, the number of trajectories, absolute throughput, latency, quality comparisons, failure rates or results from independent operators. It also says the measurements do not yet reflect Vera CPU performance for tool calling. Readers therefore cannot conclude from this source alone that every agent workload will be 30 times more efficient or that output quality will remain equivalent in all conditions.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Interactive Concept Check+10 Points
AI Agents Quiz

What is the most accurate way to describe what AI Agents can do today?

Maxaa la daawan doona xiga

The key next checks are SemiAnalysis’s review, independent testing across models and agent tasks, results that include Vera CPU tool-calling performance, and evidence that the claimed gains hold beyond peak configurations and selected trajectories.

The first point to watch is the SemiAnalysis review. It could clarify how AgentX was constructed, how many coding trajectories were measured, which settings were used, how throughput and power were calculated, and whether the comparison can be reproduced. The review may also show whether the headline result reflects a typical performance range or a best-case point. Until those details are available, NVIDIA’s figures should be treated as company-reported measurements.

Independent results across different agent tasks will be important. Coding trajectories may not represent research, customer service, data analysis or other workflows with different tool patterns and context-growth behavior. Future testing should show whether the relative advantage persists across models, request lengths, sub-agent counts, cache behavior and quality targets. The source names several models, but it does not provide a complete workload-by-workload table or establish that the claimed maximum applies broadly.

The source specifically says that the results do not yet include Vera CPU performance for tool calling. That component could be important for agents whose efficiency depends on orchestration, data access and other sequential operations outside GPU inference. Further reporting should separate GPU inference results from end-to-end agent performance and show how the Vera CPU and the other platform components affect latency, utilization and power consumption during complete workflows.

NVIDIA also claims that DSX MaxLPS can provision up to 40% more GPUs within the same megawatt budget and that NVLink 6 delivers higher packet rates and lower latency than off-the-shelf Ethernet alternatives. Those claims warrant measurements that include facility-level power behavior, cooling, networking and utilization rather than only accelerator throughput. The source also says NVFP4 increases throughput without sacrificing output quality, but it supplies no quality tests here. Independent evaluation of accuracy, reliability, latency, availability and total operating cost will determine how much of the claimed efficiency translates into practical public or commercial impact.

Tilmaamaha la xidhiidha & su'aalaha

Wakiilada AIMoodooyinka AI ayaa la sharaxayTababarka AIMustaqbalka AITijaabi waxaad taqaan - isku day kedis AI oo bilaash ahKa raadi erey AI qaamuuskeena

Cusbooneysiin iyo sixid

Sheekadan qaanuuniga ah waxaa lagu cusboonaysiiyaa meesha marka dhacdada soo koraysa ay wax iska beddesho. URLkeeda iyo taariikhda daabacaadda asalka ah weligood isma beddelaan.

  • This NVIDIA blog is a same-day primary-source elaboration of the existing Vera Rubin efficiency report. It adds details about the AgentX coding trajectories, preserved context growth, tool calls and sub-agent spawning, while also stating that the results are early, measured by NVIDIA, pending SemiAnalysis review and incomplete because Vera CPU tool-calling performance is not included.
  • This is a continuing update to NVIDIA’s reported agentic-inference efficiency results. The new post adds methodological detail about SemiAnalysis AgentX, including replayed Claude Code sessions, context reuse, tool-call timing, concurrency and user-experience metrics. It reiterates the headline claim that Vera Rubin NVL72 delivered up to 30x higher throughput per megawatt than GB300 NVL72, while explicitly noting that NVIDIA measured the result and that SemiAnalysis review is still pending.
  • This source materially advances the same Vera Rubin and Blackwell agentic-inference efficiency event already covered by the canonical update. It adds NVIDIA’s detailed claim of up to 30x higher throughput per megawatt and 35x lower token cost for Vera Rubin NVL72 versus GB300 NVL72, identifies the SemiAnalysis AgentX workload and its preserved coding-agent trajectories, describes the claimed hardware-software optimizations, and explicitly notes that the results are early and pending SemiAnalysis review and exclude Vera CPU performance for tool calling.
Eeg qoraalka sixitaanka dadweynaha
Tan faa'iido ma u heshay?