返回新聞
創新AI Understanding 簡報

NVIDIA 報告 JAX 中的無掉落 MoE 訓練速度提高了 10 倍

NVIDIA 表示,其 Transformer Engine 和 JAX 優化將 DeepSeek-V3 671B 訓練吞吐量提高了約 10 倍,並在 1,024 個 GPU 上實現了 97% 的效率。優化路徑包含在日期為 2026 年 9 月 9 日或更新的 NVIDIA NGC MaxText 容器中。

4 min readRead the primary source
Source-provided image accompanying NVIDIA reports 10x faster dropless MoE training in JAX
主要來源文件來源記錄
出版商
developer.nvidia.com
來源連結
developer.nvidia.comhttps://developer.nvidia.com/blog/accelerating-dropless-moe-training-in-jax-with-nvidia-transformer-engine/
來源類型
主要文件-我們直接閱讀的官方公告、文件、文件或第一方頁面。
背景60 秒內了解這一點

從這裡開始

關鍵術語

混合式專家 (MoE)
具有專門子網路的架構,其中每個輸入僅運行選定的專家。
記憶體(代理記憶體)
AI 代理程式跨步驟或會話使用儲存的上下文來提高連續性。
量化
將模型權重轉換為較低精確度的格式,例如 8 位元或 4 位元。
測試一下自己AI 模型解釋測驗

發生了什麼事

NVIDIA describes a set of JAX and Transformer Engine optimizations for dropless mixture-of-experts training, including grouped GEMM, fused expert-parallel dispatch and combine operations, NCCL EP, XLA multistream collectives, MXFP8 grouped , and host activation offloading. The company says an unoptimized DeepSeek-V3 training baseline reached 103 TFLOPS per GPU, compared with 1,068 TFLOPS per GPU after optimization, and that end-to-end throughput improved by about 10x. NVIDIA also reports 97% scaling efficiency at 1,024 GPUs. The optimized components are included in the NVIDIA NGC MaxText container, with instructions for testing and reproducing the configuration.

NVIDIA’s post focuses on dropless mixture-of-experts training, in which every token is processed by its selected expert even when routing creates uneven expert loads. Unlike capacity-based approaches that trim overflow or pad experts to fixed sizes, dropless training requires kernels and communication paths to handle variable token counts. NVIDIA says these irregular shapes produce ragged tensors and can leave GPUs waiting on dispatch, all-to-all communication, and expert computation.

The reported changes operate at several layers. Transformer Engine’s grouped GEMM processes variable-length expert groups in one kernel call, while fused dispatch and combine operations and the NCCL EP backend address movement of tokens between GPUs. NVIDIA also cites XLA multistream collectives, host activation offloading, and MXFP8 grouped as contributors to the result.

The company reports that its DeepSeek-V3 671B configuration increased a baseline from 103 to 1,068 TFLOPS per GPU and produced an approximately 10x end-to-end throughput gain. It also reports 97% efficiency at 1,024 GPUs. These figures are presented by NVIDIA as observations from its optimized stack, not as independently verified results.

The implementation is available to try through the NVIDIA NGC MaxText container with Transformer Engine built in. The post specifies ghcr.io/nvidia/jax:maxtext-2026-09-09 or newer and provides configuration guidance for DeepSeek-V3. It cautions that different models require different tuning. Pricing and licensing details are not provided.

來源詳情: developer.nvidia.com ↗

為什麼這很重要

Large mixture-of-experts models reduce computation by activating only selected expert networks, but their uneven token routing creates difficult communication and memory bottlenecks. If NVIDIA’s results hold beyond its reported configuration, the changes could make training very large MoE systems more efficient without dropping routed tokens or padding every expert to a fixed capacity. That matters primarily to organizations operating large NVIDIA GPU clusters, where communication overhead and wasted computation can materially affect training time and cost. The results are NVIDIA’s own claims; the source does not provide independent benchmarking or a comparison with every alternative implementation.

MoE training is increasingly important for large AI models because conditional computation can reduce the active computation per token while retaining many parameters. The engineering difficulty shifts toward routing, irregular matrix operations, memory use, and interconnect traffic. Improving those parts could reduce the time and hardware required to train models at production scale.

The practical audience is specialized: researchers and companies already using JAX, MaxText, NVIDIA GPUs, and multi-GPU or multi-node infrastructure. The source does not establish that the same gains apply to smaller systems, non-NVIDIA hardware, dense models, or workloads outside the reported DeepSeek-V3 configuration.

NVIDIA’s figures should be treated as vendor-reported performance claims. The source supplies no independent tests, detailed cost analysis, or full methodology sufficient to determine how much of the improvement comes from each optimization.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Document Size:128K tokens
Needle Placement Depth (Location in document):50% into text
Attention Context Buffer Map:
Target Fact (50%)
Equivalent Pages~320Standard book pages
Retrieval Accuracy99.9%Needle recall score
RAM / KV Cache5.1 GBMemory overhead
Prompt CachingActive~80% discount on reuse
Core takeaway: Million-token context windows allow querying whole codebases or legal archives in one prompt. However, KV cache memory scales with context length, making prompt caching crucial for real-time production.
互動式概念檢查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下來看什麼

The main question is whether the reported gains reproduce across different MoE architectures, cluster layouts, model sizes, and GPU generations. NVIDIA plans to add NVFP4, fused with GEMM, and additional all-to-all overlap. Users should also watch whether these features become standard in JAX and MaxText workflows, and whether the optimizations require substantial model-specific tuning. The source documents a container-based access path but does not state pricing, licensing terms, support commitments, or general availability beyond the cited NGC container.

Reproduction across other MoE models and hardware configurations will indicate whether the reported gains are broadly transferable or tightly coupled to DeepSeek-V3 and NVIDIA’s cluster setup.

NVIDIA says future work will add NVFP4, fused with GEMM, and additional all-to-all overlap. Those additions could further affect the performance and memory tradeoffs of the stack.

The source recommends tracking step time, TFLOPS per GPU, model FLOP utilization, grouped GEMM latency, and dispatch/combine latency. Those measurements will help separate compute gains from communication improvements.

The post does not state the container’s price, licensing conditions, support level, or whether all cited components are equally mature for production use.

相關指引和測驗

人工智慧模型解釋變形金剛人工智慧培訓AI 的未來測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?