What happened
NVIDIA submitted preview performance data for its Vera Rubin NVL72 platform to the MLPerf Inference v6.1 benchmark suite, demonstrating significant throughput improvements over the previous generation GB300 NVL72. The results indicate up to 3.7x higher throughput on the Qwen3-VL benchmark and 2.5x higher throughput on DeepSeek-R1, driven by hardware-software co-design, NVFP4 precision, and disaggregated serving techniques.
NVIDIA released preview results for the Vera Rubin NVL72 platform in the MLPerf Inference v6.1 suite, focusing on the DeepSeek-R1 and Qwen3-VL benchmarks. The company reported that Vera Rubin NVL72 delivers up to 3.7x higher throughput than the GB300 NVL72 on Qwen3-VL across offline, server, and interactive scenarios when using vLLM with the NVIDIA Dynamo open-source inference framework. For the DeepSeek-R1 benchmark, using the NVIDIA TensorRT-LLM library, throughput was up to 2.5x higher than the GB300 NVL72.
The performance gains are attributed to full-stack co-design, including enhanced Tensor Cores, the Transformer Engine, and NVFP4 precision, which reduces memory footprint for model weights, attention, and KV cache. The submissions heavily utilized disaggregated serving, separating prefill and decode stages, along with large-scale expert parallelism for mixture-of-experts layers. The NVL72 scale-up domain, powered by sixth-generation NVLink and NVLink Switch, provided the interconnect foundation necessary for these techniques at rack scale.
NVIDIA also highlighted scaling efficiency, noting that the DeepSeek-R1 submission scaled from a single GB300 NVL72 rack to four racks (288 GPUs) with 99% scaling efficiency in the offline scenario. In agentic workloads, specifically the SemiAnalysis AgentX benchmark, Vera Rubin NVL72 delivered 30x better performance than GB300 NVL72 in preview testing. Partner Nebius also submitted Vera Rubin NVL72 preview results, demonstrating similar performance characteristics.
The release includes results from the broader NVIDIA ecosystem, with 19 partners submitting data, including ASUS, Azure, Cisco, CoreWeave, and Oracle Cloud Infrastructure. NVIDIA also submitted Jetson AGX Thor results using TensorRT Edge-LLM on the new Edge-Agentic benchmark with Qwen3.6-27B. Post-submission results for GPT-OSS-120B and DLRMv3 showed further gains but are not yet verified by MLCommons.
Source details: blogs.nvidia.com ↗
Why it matters
These results provide concrete evidence of the performance trajectory for next-generation AI inference infrastructure, directly impacting the cost-per-token economics for organizations deploying large language models. By demonstrating near-linear scaling efficiency across multiple racks and substantial gains in agentic workloads, the data helps enterprises forecast infrastructure requirements and validate the economic viability of scaling AI deployments. This matters because inference costs are a primary driver of AI product profitability and accessibility.
The primary significance of these results lies in the quantification of inference economics. By demonstrating that each Vera Rubin NVL72 rack can generate significantly more tokens and serve more users than a GB300 NVL72 rack, NVIDIA provides a clear metric for cost-per-token reduction. This is critical for organizations where inference costs constitute a major portion of operational expenses, as it directly translates to higher revenue potential or lower service costs for the same workload.
The emphasis on scaling efficiency addresses a common pain point in AI infrastructure: the non-linear relationship between hardware addition and throughput gains. Achieving 99% scaling efficiency across 288 GPUs suggests that the architecture and interconnects are effectively mitigating communication bottlenecks, allowing organizations to predict performance gains more accurately when expanding their AI factories.
The inclusion of agentic workload benchmarks, such as SemiAnalysis AgentX, signals a shift in how AI performance is measured. As AI systems move from single-turn responses to multi-step reasoning and action, traditional throughput metrics may not fully capture utility. The 30x performance improvement in this specific domain indicates that next-generation hardware is specifically optimized for the latency and compute demands of agentic AI, which is a growing sector in enterprise applications.
What to watch next
Monitor the final verification of these results by MLCommons and the subsequent release of the Vera Rubin platform to the market. Additionally, track the adoption of the new MLPerf Endpoints benchmark for agentic inference, which will standardize performance metrics for multi-step AI agents, and observe how competitors respond to these specific throughput and scaling efficiency benchmarks.
The final verification of these preview results by MLCommons is the immediate next step. Until verified, these numbers are preliminary and subject to change. The official release will provide the definitive performance baseline for the Vera Rubin platform.
The market availability and pricing of the Vera Rubin NVL72 platform will determine its practical impact. While the performance gains are documented, the cost of the hardware and the transition period from GB300 will influence adoption rates. Organizations will need to weigh the performance benefits against the capital expenditure required for the upgrade.
The development and adoption of the MLPerf Endpoints benchmark for agentic inference will be crucial. As this benchmark becomes standardized, it will provide a more comprehensive view of AI system capabilities beyond raw token throughput, potentially reshaping how vendors market and compare their platforms.