Komawa Labarai
Bidi'aAI Understanding takaitaccen bayani

NVIDIA Vera Rubin NVL72 abubuwan da ke jagorantar MLPerf Inference v6.1 sakamako

NVIDIA ya fitar da sakamakon samfoti wanda ke nuna Vera Rubin NVL72 ya cimma har zuwa 3.7x mafi girma kayan aiki fiye da GB300 NVL72 akan Qwen3-VL da 2.5x akan DeepSeek-R1 a cikin MLperf Inference v6.1 suite.

4 min readRead the primary source
Source-provided image accompanying NVIDIA Vera Rubin NVL72 posts leading MLPerf Inference v6.1 results
Takardun tushe na farkoAn rubuta tushen tushe
Mawallafi
blogs.nvidia.com
Tushen hanyar haɗin gwiwa
blogs.nvidia.comhttps://blogs.nvidia.com/blog/vera-rubin-nvl72-mlperf-inference/
Nau'in tushe
Takardun farko - sanarwar hukuma, takarda, yin rajista, ko shafi na farko da muka karanta kai tsaye.
MaganaFahimtar wannan a cikin daƙiƙa 60

Fara a nan

Mabuɗin sharuddan

Inference
Lokaci lokacin aiki inda ƙwararren ƙirar ke haifar da tsinkaya ko fitarwa.
Babban Samfurin Harshe (LLM)
Samfurin harshe da aka horar akan babban haɗin gwiwar rubutu don samarwa da tantance rubutu.
Ƙwaƙwalwar ajiya (Agent Memory)
Mahallin da aka adana wani wakilin AI yana amfani da matakai ko zaman don inganta ci gaba.
Gwada kankaAI Model An Bayyana Tambayoyi

Me ya faru

NVIDIA submitted preview performance data for its Vera Rubin NVL72 platform to the MLPerf v6.1 benchmark suite, demonstrating significant throughput improvements over the previous generation GB300 NVL72. The results indicate up to 3.7x higher throughput on the Qwen3-VL benchmark and 2.5x higher throughput on DeepSeek-R1, driven by hardware-software co-design, NVFP4 precision, and disaggregated serving techniques.

NVIDIA released preview results for the Vera Rubin NVL72 platform in the MLPerf v6.1 suite, focusing on the DeepSeek-R1 and Qwen3-VL benchmarks. The company reported that Vera Rubin NVL72 delivers up to 3.7x higher throughput than the GB300 NVL72 on Qwen3-VL across offline, server, and interactive scenarios when using vLLM with the NVIDIA Dynamo open-source inference framework. For the DeepSeek-R1 benchmark, using the NVIDIA TensorRT-LLM library, throughput was up to 2.5x higher than the GB300 NVL72.

The performance gains are attributed to full-stack co-design, including enhanced Tensor Cores, the Transformer Engine, and NVFP4 precision, which reduces memory footprint for model weights, attention, and KV cache. The submissions heavily utilized disaggregated serving, separating prefill and decode stages, along with large-scale expert parallelism for mixture-of-experts layers. The NVL72 scale-up domain, powered by sixth-generation NVLink and NVLink Switch, provided the interconnect foundation necessary for these techniques at rack scale.

NVIDIA also highlighted scaling efficiency, noting that the DeepSeek-R1 submission scaled from a single GB300 NVL72 rack to four racks (288 GPUs) with 99% scaling efficiency in the offline scenario. In agentic workloads, specifically the SemiAnalysis AgentX benchmark, Vera Rubin NVL72 delivered 30x better performance than GB300 NVL72 in preview testing. Partner Nebius also submitted Vera Rubin NVL72 preview results, demonstrating similar performance characteristics.

The release includes results from the broader NVIDIA ecosystem, with 19 partners submitting data, including ASUS, Azure, Cisco, CoreWeave, and Oracle Cloud Infrastructure. NVIDIA also submitted Jetson AGX Thor results using TensorRT Edge-LLM on the new Edge-Agentic benchmark with Qwen3.6-27B. Post-submission results for GPT-OSS-120B and DLRMv3 showed further gains but are not yet verified by MLCommons.

Bayanan tushe: blogs.nvidia.com ↗

Me ya sa yake da mahimmanci

These results provide concrete evidence of the performance trajectory for next-generation AI infrastructure, directly impacting the cost-per-token economics for organizations deploying large language models. By demonstrating near-linear scaling efficiency across multiple racks and substantial gains in agentic workloads, the data helps enterprises forecast infrastructure requirements and validate the economic viability of scaling AI deployments. This matters because inference costs are a primary driver of AI product profitability and accessibility.

The primary significance of these results lies in the quantification of economics. By demonstrating that each Vera Rubin NVL72 rack can generate significantly more tokens and serve more users than a GB300 NVL72 rack, NVIDIA provides a clear metric for cost-per-token reduction. This is critical for organizations where inference costs constitute a major portion of operational expenses, as it directly translates to higher revenue potential or lower service costs for the same workload.

The emphasis on scaling efficiency addresses a common pain point in AI infrastructure: the non-linear relationship between hardware addition and throughput gains. Achieving 99% scaling efficiency across 288 GPUs suggests that the architecture and interconnects are effectively mitigating communication bottlenecks, allowing organizations to predict performance gains more accurately when expanding their AI factories.

The inclusion of agentic workload benchmarks, such as SemiAnalysis AgentX, signals a shift in how AI performance is measured. As AI systems move from single-turn responses to multi-step reasoning and action, traditional throughput metrics may not fully capture utility. The 30x performance improvement in this specific domain indicates that next-generation hardware is specifically optimized for the latency and compute demands of agentic AI, which is a growing sector in enterprise applications.

Interactive Mechanism

Ingantacciyar hanyar sadarwa: Yadda A zahiri yake Aiki

Bincika fasahar da ke bayan wannan ci gaban ta hanyar mu'amala.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Duba ra'ayi na hulɗa+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Abin kallo na gaba

Monitor the final verification of these results by MLCommons and the subsequent release of the Vera Rubin platform to the market. Additionally, track the adoption of the new MLPerf Endpoints benchmark for agentic , which will standardize performance metrics for multi-step AI agents, and observe how competitors respond to these specific throughput and scaling efficiency benchmarks.

The final verification of these preview results by MLCommons is the immediate next step. Until verified, these numbers are preliminary and subject to change. The official release will provide the definitive performance baseline for the Vera Rubin platform.

The market availability and pricing of the Vera Rubin NVL72 platform will determine its practical impact. While the performance gains are documented, the cost of the hardware and the transition period from GB300 will influence adoption rates. Organizations will need to weigh the performance benefits against the capital expenditure required for the upgrade.

The development and adoption of the MLPerf Endpoints benchmark for agentic will be crucial. As this benchmark becomes standardized, it will provide a more comprehensive view of AI system capabilities beyond raw token throughput, potentially reshaping how vendors market and compare their platforms.

Jagorori masu alaƙa & tambayoyin tambayoyi

AI Model ya bayyanaAI horoMakomar AIGwada abin da kuka sani - gwada gwajin AI kyautaNemo kalmar AI a cikin ƙamus ɗin muBi samfurin AI na sakin tracker
An sami wannan yana da amfani?