Apa yang berlaku
NVIDIA reports new performance measurements for Vera Rubin NVL72 systems running agentic-AI workloads, claiming substantially higher throughput per megawatt and lower token costs than its GB300 NVL72 platform.
NVIDIA said on August 24 that Vera Rubin NVL72 systems delivered up to 30 times higher agentic-AI inference throughput per megawatt than NVIDIA GB300 NVL72 systems. The company also claimed up to 35 times lower cost per million tokens. These are NVIDIA’s figures, not independently established results, and the source presents them as early performance data rather than a general guarantee for every workload. The comparison is specifically between NVIDIA’s Vera Rubin and GB300 NVL72 systems.
The measurements used SemiAnalysis AgentX, which NVIDIA describes as a workload built from recorded real-world agentic coding sessions. According to the source, the preserves context growth, tool calls and the spawning of sub-agents. NVIDIA says the workload was intended to measure an entire agent workflow rather than a single request. The source lists Kimi K3, MiniMax M3, GLM5.3, Qwen3.5 and DeepSeek V4 Pro among the models used in the broader AgentX results, and gives DeepSeek V4 Pro as an example where Vera Rubin reached as much as 30 times the throughput per megawatt of GB300 NVL72.
NVIDIA frames agentic workloads as unusually demanding because an agent may repeatedly retrieve information, call tools, invoke other agents and carry accumulated context into later steps. The source contrasts these sessions with chat or document-summarization requests that typically involve input and output sequences of 1,000 to 8,000 tokens. It says agentic sessions can reach hundreds of thousands of input tokens, with substantial variation in input and output lengths. NVIDIA also cites OpenRouter data saying agentic workloads consume 15 times more tokens than a simple chat request, though the source does not provide the underlying data or methodology.
The company attributes the reported gains to a rack-scale platform and software stack that combine disaggregated prefill and decode serving, rate matching, expert parallelism, distributed key-value caching, cache-aware routing, fused CUDA kernels, Tensor Cores, Engine, NVFP4 and high-bandwidth NVLink interconnects. NVIDIA says the full Vera Rubin platform includes seven chip types, including the Vera CPU, Groq 3 LPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX and ConnectX-9 SuperNIC. The source says Vera Rubin is in full production and scaling across its ecosystem, but it does not specify customer deployments, volumes, prices or general availability terms.
Butiran sumber: blogs.nvidia.com ↗
Mengapa ia penting
If independently confirmed across varied workloads, the claimed efficiency gains could affect the energy, operating-cost and infrastructure requirements of large-scale AI agents. The evidence is currently limited to NVIDIA’s measurements and a workload review that is still pending.
The central significance of the announcement is energy efficiency. NVIDIA argues that throughput per megawatt can determine how much agentic work an AI factory performs within a fixed power budget. If the reported comparison survives independent review, operators with constrained electricity capacity could potentially run more agent sessions without increasing their power allocation. That could matter as agents perform longer sequences of retrieval, reasoning and tool use than ordinary chat systems. The source does not establish how the result would translate into savings at a particular facility.
The claimed token-cost reduction matters because agent systems may generate and process much more context than a single-turn application. NVIDIA says cost per million tokens affects the profit margin of AI-factory operators. Lower inference cost could make some long-running coding, research or customer-service workflows easier to operate continuously. That conclusion remains conditional: the source gives relative results, not an absolute cost per token, a full system price, deployment cost, electricity price or total cost of ownership.
The measurement approach is notable because it attempts to represent agent behavior over multiple steps. A based only on one prompt and one response could miss the overhead created by repeated tool calls, growing context and coordination among sub-agents. NVIDIA says AgentX retains those features from recorded coding sessions. That makes the result more relevant to the type of workloads the company is targeting, but it also means the outcome may depend heavily on the selected trajectories, models, context lengths, serving configuration and software optimizations.
There are material limits to the evidence. The results were measured by NVIDIA, the page says they are pending SemiAnalysis review, and the headline figures use “up to,” which describes a maximum rather than an average. The source does not report confidence intervals, the number of trajectories, absolute throughput, latency, quality comparisons, failure rates or results from independent operators. It also says the measurements do not yet reflect Vera CPU performance for tool calling. Readers therefore cannot conclude from this source alone that every agent workload will be 30 times more efficient or that output quality will remain equivalent in all conditions.
Interactive Mechanism: How It Actually Works
Explore the underlying technology behind this development interactively.
What is the most accurate way to describe what AI Agents can do today?
Apa yang perlu ditonton seterusnya
The key next checks are SemiAnalysis’s review, independent testing across models and agent tasks, results that include Vera CPU tool-calling performance, and evidence that the claimed gains hold beyond peak configurations and selected trajectories.
The first point to watch is the SemiAnalysis review. It could clarify how AgentX was constructed, how many coding trajectories were measured, which settings were used, how throughput and power were calculated, and whether the comparison can be reproduced. The review may also show whether the headline result reflects a typical performance range or a best-case point. Until those details are available, NVIDIA’s figures should be treated as company-reported measurements.
Independent results across different agent tasks will be important. Coding trajectories may not represent research, customer service, data analysis or other workflows with different tool patterns and context-growth behavior. Future testing should show whether the relative advantage persists across models, request lengths, sub-agent counts, cache behavior and quality targets. The source names several models, but it does not provide a complete workload-by-workload table or establish that the claimed maximum applies broadly.
The source specifically says that the results do not yet include Vera CPU performance for tool calling. That component could be important for agents whose efficiency depends on orchestration, data access and other sequential operations outside GPU inference. Further reporting should separate GPU inference results from end-to-end agent performance and show how the Vera CPU and the other platform components affect latency, utilization and power consumption during complete workflows.
NVIDIA also claims that DSX MaxLPS can provision up to 40% more GPUs within the same megawatt budget and that NVLink 6 delivers higher packet rates and lower latency than off-the-shelf Ethernet alternatives. Those claims warrant measurements that include facility-level power behavior, cooling, networking and utilization rather than only accelerator throughput. The source also says NVFP4 increases throughput without sacrificing output quality, but it supplies no quality tests here. Independent evaluation of accuracy, reliability, latency, availability and total operating cost will determine how much of the claimed efficiency translates into practical public or commercial impact.