Torna alle notizie
InnovazioneAI Understanding briefing

NVIDIA Vera Rubin NVL72 pubblica i principali risultati di MLPerf Inference v6.1

NVIDIA ha pubblicato i risultati di anteprima che mostrano che Vera Rubin NVL72 raggiunge un throughput fino a 3,7 volte superiore rispetto a GB300 NVL72 su Qwen3-VL e 2,5 volte su DeepSeek-R1 nella suite MLPerf Inference v6.1.

4 min readRead the primary source
Source-provided image accompanying NVIDIA Vera Rubin NVL72 posts leading MLPerf Inference v6.1 results
Documento di origine primariaFonte registrata
Editore
blogs.nvidia.com
Collegamento alla fonte
blogs.nvidia.comhttps://blogs.nvidia.com/blog/vera-rubin-nvl72-mlperf-inference/
Tipo di fonte
Documento principale: un annuncio ufficiale, un documento, un documento o una pagina proprietaria che leggiamo direttamente.
ContestoComprendilo in 60 secondi

Inizia qui

Termini chiave

Inferenza
La fase di runtime in cui un modello addestrato genera previsioni o output.
Modello linguistico di grandi dimensioni (LLM)
Un modello linguistico addestrato su enormi corpora di testo per generare e analizzare testo.
Memoria (memoria dell'agente)
Contesto archiviato che un agente AI utilizza attraverso passaggi o sessioni per migliorare la continuità.
Mettiti alla provaQuiz sulla spiegazione dei modelli di intelligenza artificiale

Cosa è successo

NVIDIA submitted preview performance data for its Vera Rubin NVL72 platform to the MLPerf v6.1 benchmark suite, demonstrating significant throughput improvements over the previous generation GB300 NVL72. The results indicate up to 3.7x higher throughput on the Qwen3-VL benchmark and 2.5x higher throughput on DeepSeek-R1, driven by hardware-software co-design, NVFP4 precision, and disaggregated serving techniques.

NVIDIA released preview results for the Vera Rubin NVL72 platform in the MLPerf v6.1 suite, focusing on the DeepSeek-R1 and Qwen3-VL benchmarks. The company reported that Vera Rubin NVL72 delivers up to 3.7x higher throughput than the GB300 NVL72 on Qwen3-VL across offline, server, and interactive scenarios when using vLLM with the NVIDIA Dynamo open-source inference framework. For the DeepSeek-R1 benchmark, using the NVIDIA TensorRT-LLM library, throughput was up to 2.5x higher than the GB300 NVL72.

The performance gains are attributed to full-stack co-design, including enhanced Tensor Cores, the Transformer Engine, and NVFP4 precision, which reduces memory footprint for model weights, attention, and KV cache. The submissions heavily utilized disaggregated serving, separating prefill and decode stages, along with large-scale expert parallelism for mixture-of-experts layers. The NVL72 scale-up domain, powered by sixth-generation NVLink and NVLink Switch, provided the interconnect foundation necessary for these techniques at rack scale.

NVIDIA also highlighted scaling efficiency, noting that the DeepSeek-R1 submission scaled from a single GB300 NVL72 rack to four racks (288 GPUs) with 99% scaling efficiency in the offline scenario. In agentic workloads, specifically the SemiAnalysis AgentX benchmark, Vera Rubin NVL72 delivered 30x better performance than GB300 NVL72 in preview testing. Partner Nebius also submitted Vera Rubin NVL72 preview results, demonstrating similar performance characteristics.

The release includes results from the broader NVIDIA ecosystem, with 19 partners submitting data, including ASUS, Azure, Cisco, CoreWeave, and Oracle Cloud Infrastructure. NVIDIA also submitted Jetson AGX Thor results using TensorRT Edge-LLM on the new Edge-Agentic benchmark with Qwen3.6-27B. Post-submission results for GPT-OSS-120B and DLRMv3 showed further gains but are not yet verified by MLCommons.

Dettagli della fonte: blogs.nvidia.com ↗

Perché è importante

These results provide concrete evidence of the performance trajectory for next-generation AI infrastructure, directly impacting the cost-per-token economics for organizations deploying large language models. By demonstrating near-linear scaling efficiency across multiple racks and substantial gains in agentic workloads, the data helps enterprises forecast infrastructure requirements and validate the economic viability of scaling AI deployments. This matters because inference costs are a primary driver of AI product profitability and accessibility.

The primary significance of these results lies in the quantification of economics. By demonstrating that each Vera Rubin NVL72 rack can generate significantly more tokens and serve more users than a GB300 NVL72 rack, NVIDIA provides a clear metric for cost-per-token reduction. This is critical for organizations where inference costs constitute a major portion of operational expenses, as it directly translates to higher revenue potential or lower service costs for the same workload.

The emphasis on scaling efficiency addresses a common pain point in AI infrastructure: the non-linear relationship between hardware addition and throughput gains. Achieving 99% scaling efficiency across 288 GPUs suggests that the architecture and interconnects are effectively mitigating communication bottlenecks, allowing organizations to predict performance gains more accurately when expanding their AI factories.

The inclusion of agentic workload benchmarks, such as SemiAnalysis AgentX, signals a shift in how AI performance is measured. As AI systems move from single-turn responses to multi-step reasoning and action, traditional throughput metrics may not fully capture utility. The 30x performance improvement in this specific domain indicates that next-generation hardware is specifically optimized for the latency and compute demands of agentic AI, which is a growing sector in enterprise applications.

Interactive Mechanism

Meccanismo interattivo: come funziona realmente

Esplora la tecnologia alla base di questo sviluppo in modo interattivo.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Verifica concettuale interattiva+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Cosa guardare dopo

Monitor the final verification of these results by MLCommons and the subsequent release of the Vera Rubin platform to the market. Additionally, track the adoption of the new MLPerf Endpoints benchmark for agentic , which will standardize performance metrics for multi-step AI agents, and observe how competitors respond to these specific throughput and scaling efficiency benchmarks.

The final verification of these preview results by MLCommons is the immediate next step. Until verified, these numbers are preliminary and subject to change. The official release will provide the definitive performance baseline for the Vera Rubin platform.

The market availability and pricing of the Vera Rubin NVL72 platform will determine its practical impact. While the performance gains are documented, the cost of the hardware and the transition period from GB300 will influence adoption rates. Organizations will need to weigh the performance benefits against the capital expenditure required for the upgrade.

The development and adoption of the MLPerf Endpoints benchmark for agentic will be crucial. As this benchmark becomes standardized, it will provide a more comprehensive view of AI system capabilities beyond raw token throughput, potentially reshaping how vendors market and compare their platforms.

Guide e quiz correlati

Spiegazione dei modelli di intelligenza artificialeFormazione sull'intelligenza artificialeFuturo dell'IAMetti alla prova ciò che sai: prova un quiz gratuito sull'intelligenza artificialeCerca un termine AI nel nostro glossarioSegui il tracker del rilascio del modello AI
Lo hai trovato utile?