Retour aux Actualités
InnovationBriefing AI Understanding

NVIDIA Vera Rubin NVL72 publie les principaux résultats de MLPerf Inference v6.1

NVIDIA a publié des résultats préliminaires montrant que Vera Rubin NVL72 atteint un débit jusqu'à 3,7 fois supérieur à celui du GB300 NVL72 sur Qwen3-VL et 2,5x sur DeepSeek-R1 dans la suite MLPerf Inference v6.1.

4 min readRead the primary source
Source-provided image accompanying NVIDIA Vera Rubin NVL72 posts leading MLPerf Inference v6.1 results
Document de source principaleSource enregistrée
Éditeur
blogs.nvidia.com
Lien source
blogs.nvidia.comhttps://blogs.nvidia.com/blog/vera-rubin-nvl72-mlperf-inference/
Type de source
Document principal : une annonce officielle, un document, un dépôt ou une page de première partie que nous lisons directement.
ContexteComprenez cela en 60 secondes

Commencez ici

Termes clés

Inférence
Phase d'exécution au cours de laquelle un modèle entraîné génère des prédictions ou des sorties.
Grand modèle linguistique (LLM)
Un modèle de langage formé sur des corpus de textes massifs pour générer et analyser du texte.
Mémoire (mémoire de l'agent)
Contexte stocké qu'un agent IA utilise au fil des étapes ou des sessions pour améliorer la continuité.
Testez-vousQuiz sur les modèles d'IA expliqués

Que s'est-il passé

NVIDIA submitted preview performance data for its Vera Rubin NVL72 platform to the MLPerf v6.1 benchmark suite, demonstrating significant throughput improvements over the previous generation GB300 NVL72. The results indicate up to 3.7x higher throughput on the Qwen3-VL benchmark and 2.5x higher throughput on DeepSeek-R1, driven by hardware-software co-design, NVFP4 precision, and disaggregated serving techniques.

NVIDIA released preview results for the Vera Rubin NVL72 platform in the MLPerf v6.1 suite, focusing on the DeepSeek-R1 and Qwen3-VL benchmarks. The company reported that Vera Rubin NVL72 delivers up to 3.7x higher throughput than the GB300 NVL72 on Qwen3-VL across offline, server, and interactive scenarios when using vLLM with the NVIDIA Dynamo open-source inference framework. For the DeepSeek-R1 benchmark, using the NVIDIA TensorRT-LLM library, throughput was up to 2.5x higher than the GB300 NVL72.

The performance gains are attributed to full-stack co-design, including enhanced Tensor Cores, the Transformer Engine, and NVFP4 precision, which reduces memory footprint for model weights, attention, and KV cache. The submissions heavily utilized disaggregated serving, separating prefill and decode stages, along with large-scale expert parallelism for mixture-of-experts layers. The NVL72 scale-up domain, powered by sixth-generation NVLink and NVLink Switch, provided the interconnect foundation necessary for these techniques at rack scale.

NVIDIA also highlighted scaling efficiency, noting that the DeepSeek-R1 submission scaled from a single GB300 NVL72 rack to four racks (288 GPUs) with 99% scaling efficiency in the offline scenario. In agentic workloads, specifically the SemiAnalysis AgentX benchmark, Vera Rubin NVL72 delivered 30x better performance than GB300 NVL72 in preview testing. Partner Nebius also submitted Vera Rubin NVL72 preview results, demonstrating similar performance characteristics.

The release includes results from the broader NVIDIA ecosystem, with 19 partners submitting data, including ASUS, Azure, Cisco, CoreWeave, and Oracle Cloud Infrastructure. NVIDIA also submitted Jetson AGX Thor results using TensorRT Edge-LLM on the new Edge-Agentic benchmark with Qwen3.6-27B. Post-submission results for GPT-OSS-120B and DLRMv3 showed further gains but are not yet verified by MLCommons.

Détails de la source: blogs.nvidia.com ↗

Pourquoi c'est important

These results provide concrete evidence of the performance trajectory for next-generation AI infrastructure, directly impacting the cost-per-token economics for organizations deploying large language models. By demonstrating near-linear scaling efficiency across multiple racks and substantial gains in agentic workloads, the data helps enterprises forecast infrastructure requirements and validate the economic viability of scaling AI deployments. This matters because inference costs are a primary driver of AI product profitability and accessibility.

The primary significance of these results lies in the quantification of economics. By demonstrating that each Vera Rubin NVL72 rack can generate significantly more tokens and serve more users than a GB300 NVL72 rack, NVIDIA provides a clear metric for cost-per-token reduction. This is critical for organizations where inference costs constitute a major portion of operational expenses, as it directly translates to higher revenue potential or lower service costs for the same workload.

The emphasis on scaling efficiency addresses a common pain point in AI infrastructure: the non-linear relationship between hardware addition and throughput gains. Achieving 99% scaling efficiency across 288 GPUs suggests that the architecture and interconnects are effectively mitigating communication bottlenecks, allowing organizations to predict performance gains more accurately when expanding their AI factories.

The inclusion of agentic workload benchmarks, such as SemiAnalysis AgentX, signals a shift in how AI performance is measured. As AI systems move from single-turn responses to multi-step reasoning and action, traditional throughput metrics may not fully capture utility. The 30x performance improvement in this specific domain indicates that next-generation hardware is specifically optimized for the latency and compute demands of agentic AI, which is a growing sector in enterprise applications.

Interactive Mechanism

Mécanisme interactif : comment cela fonctionne réellement

Explorez de manière interactive la technologie sous-jacente à ce développement.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Vérification de concept interactive+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Que regarder ensuite

Monitor the final verification of these results by MLCommons and the subsequent release of the Vera Rubin platform to the market. Additionally, track the adoption of the new MLPerf Endpoints benchmark for agentic , which will standardize performance metrics for multi-step AI agents, and observe how competitors respond to these specific throughput and scaling efficiency benchmarks.

The final verification of these preview results by MLCommons is the immediate next step. Until verified, these numbers are preliminary and subject to change. The official release will provide the definitive performance baseline for the Vera Rubin platform.

The market availability and pricing of the Vera Rubin NVL72 platform will determine its practical impact. While the performance gains are documented, the cost of the hardware and the transition period from GB300 will influence adoption rates. Organizations will need to weigh the performance benefits against the capital expenditure required for the upgrade.

The development and adoption of the MLPerf Endpoints benchmark for agentic will be crucial. As this benchmark becomes standardized, it will provide a more comprehensive view of AI system capabilities beyond raw token throughput, potentially reshaping how vendors market and compare their platforms.

Guides et quiz associés

Modèles d'IA expliquésFormation IAAvenir de l'IATestez ce que vous savez : essayez un quiz gratuit sur l'IARecherchez un terme d'IA dans notre glossaireSuivez le suivi des versions du modèle AI
Vous avez trouvé cela utile ?