뉴스로 돌아가기
혁신AI Understanding 브리핑

NVIDIA Vera Rubin NVL72, 최고의 MLPerf 추론 v6.1 결과 게시

NVIDIA는 Vera Rubin NVL72가 MLPerf Inference v6.1 제품군의 Qwen3-VL에서 GB300 NVL72보다 최대 3.7배, DeepSeek-R1에서 2.5배 더 높은 처리량을 달성했음을 보여주는 미리 보기 결과를 발표했습니다.

4 min readRead the primary source
Source-provided image accompanying NVIDIA Vera Rubin NVL72 posts leading MLPerf Inference v6.1 results
기본 소스 문서녹음된 소스
출판사
blogs.nvidia.com
소스 링크
blogs.nvidia.comhttps://blogs.nvidia.com/blog/vera-rubin-nvl72-mlperf-inference/
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

추론
훈련된 모델이 예측 또는 출력을 생성하는 런타임 단계입니다.
대형 언어 모델(LLM)
텍스트를 생성하고 분석하기 위해 대규모 텍스트 말뭉치를 학습한 언어 모델입니다.
메모리(에이전트 메모리)
AI 에이전트는 연속성을 향상하기 위해 여러 단계 또는 세션에서 사용하는 저장된 컨텍스트입니다.
자신을 테스트해 보세요AI 모델 설명 퀴즈

무슨 일이 일어났나요?

NVIDIA submitted preview performance data for its Vera Rubin NVL72 platform to the MLPerf v6.1 benchmark suite, demonstrating significant throughput improvements over the previous generation GB300 NVL72. The results indicate up to 3.7x higher throughput on the Qwen3-VL benchmark and 2.5x higher throughput on DeepSeek-R1, driven by hardware-software co-design, NVFP4 precision, and disaggregated serving techniques.

NVIDIA released preview results for the Vera Rubin NVL72 platform in the MLPerf v6.1 suite, focusing on the DeepSeek-R1 and Qwen3-VL benchmarks. The company reported that Vera Rubin NVL72 delivers up to 3.7x higher throughput than the GB300 NVL72 on Qwen3-VL across offline, server, and interactive scenarios when using vLLM with the NVIDIA Dynamo open-source inference framework. For the DeepSeek-R1 benchmark, using the NVIDIA TensorRT-LLM library, throughput was up to 2.5x higher than the GB300 NVL72.

The performance gains are attributed to full-stack co-design, including enhanced Tensor Cores, the Transformer Engine, and NVFP4 precision, which reduces memory footprint for model weights, attention, and KV cache. The submissions heavily utilized disaggregated serving, separating prefill and decode stages, along with large-scale expert parallelism for mixture-of-experts layers. The NVL72 scale-up domain, powered by sixth-generation NVLink and NVLink Switch, provided the interconnect foundation necessary for these techniques at rack scale.

NVIDIA also highlighted scaling efficiency, noting that the DeepSeek-R1 submission scaled from a single GB300 NVL72 rack to four racks (288 GPUs) with 99% scaling efficiency in the offline scenario. In agentic workloads, specifically the SemiAnalysis AgentX benchmark, Vera Rubin NVL72 delivered 30x better performance than GB300 NVL72 in preview testing. Partner Nebius also submitted Vera Rubin NVL72 preview results, demonstrating similar performance characteristics.

The release includes results from the broader NVIDIA ecosystem, with 19 partners submitting data, including ASUS, Azure, Cisco, CoreWeave, and Oracle Cloud Infrastructure. NVIDIA also submitted Jetson AGX Thor results using TensorRT Edge-LLM on the new Edge-Agentic benchmark with Qwen3.6-27B. Post-submission results for GPT-OSS-120B and DLRMv3 showed further gains but are not yet verified by MLCommons.

소스 세부정보: blogs.nvidia.com ↗

왜 중요한가요?

These results provide concrete evidence of the performance trajectory for next-generation AI infrastructure, directly impacting the cost-per-token economics for organizations deploying large language models. By demonstrating near-linear scaling efficiency across multiple racks and substantial gains in agentic workloads, the data helps enterprises forecast infrastructure requirements and validate the economic viability of scaling AI deployments. This matters because inference costs are a primary driver of AI product profitability and accessibility.

The primary significance of these results lies in the quantification of economics. By demonstrating that each Vera Rubin NVL72 rack can generate significantly more tokens and serve more users than a GB300 NVL72 rack, NVIDIA provides a clear metric for cost-per-token reduction. This is critical for organizations where inference costs constitute a major portion of operational expenses, as it directly translates to higher revenue potential or lower service costs for the same workload.

The emphasis on scaling efficiency addresses a common pain point in AI infrastructure: the non-linear relationship between hardware addition and throughput gains. Achieving 99% scaling efficiency across 288 GPUs suggests that the architecture and interconnects are effectively mitigating communication bottlenecks, allowing organizations to predict performance gains more accurately when expanding their AI factories.

The inclusion of agentic workload benchmarks, such as SemiAnalysis AgentX, signals a shift in how AI performance is measured. As AI systems move from single-turn responses to multi-step reasoning and action, traditional throughput metrics may not fully capture utility. The 30x performance improvement in this specific domain indicates that next-generation hardware is specifically optimized for the latency and compute demands of agentic AI, which is a growing sector in enterprise applications.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
대화형 개념 확인+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

다음에 무엇을 볼 것인가

Monitor the final verification of these results by MLCommons and the subsequent release of the Vera Rubin platform to the market. Additionally, track the adoption of the new MLPerf Endpoints benchmark for agentic , which will standardize performance metrics for multi-step AI agents, and observe how competitors respond to these specific throughput and scaling efficiency benchmarks.

The final verification of these preview results by MLCommons is the immediate next step. Until verified, these numbers are preliminary and subject to change. The official release will provide the definitive performance baseline for the Vera Rubin platform.

The market availability and pricing of the Vera Rubin NVL72 platform will determine its practical impact. While the performance gains are documented, the cost of the hardware and the transition period from GB300 will influence adoption rates. Organizations will need to weigh the performance benefits against the capital expenditure required for the upgrade.

The development and adoption of the MLPerf Endpoints benchmark for agentic will be crucial. As this benchmark becomes standardized, it will provide a more comprehensive view of AI system capabilities beyond raw token throughput, potentially reshaping how vendors market and compare their platforms.

관련 가이드 및 퀴즈

AI 모델 설명AI 트레이닝AI의 미래알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.
이것이 유용하다고 생각하시나요?