Înapoi la Știri
InovațieAI Understanding briefing

NVIDIA raportează antrenament MoE fără picături de 10 ori mai rapid în JAX

NVIDIA spune că optimizările sale Transformer Engine și JAX au crescut randamentul de antrenament DeepSeek-V3 671B de aproximativ 10 ori și au atins o eficiență de 97% pe 1.024 de GPU-uri. Calea optimizată este inclusă într-un container NVIDIA NGC MaxText datat 9 septembrie 2026 sau mai nou.

4 min readRead the primary source
Source-provided image accompanying NVIDIA reports 10x faster dropless MoE training in JAX
Document sursă primarăSursa înregistrată
Editor
developer.nvidia.com
Link sursă
developer.nvidia.comhttps://developer.nvidia.com/blog/accelerating-dropless-moe-training-in-jax-with-nvidia-transformer-engine/
Tip sursă
Document principal — un anunț oficial, hârtie, depunere sau pagină primară pe care o citim direct.
ContextÎnțelege asta în 60 de secunde

Începeți de aici

Termeni cheie

Amestec de experți (MoE)
O arhitectură cu subrețele specializate în care doar experții selectați rulează per intrare.
Memorie (Memorie agent)
Context stocat pe care un agent AI îl folosește în pași sau sesiuni pentru a îmbunătăți continuitatea.
Cuantizarea
Conversia greutăților modelului în formate de precizie mai scăzută, cum ar fi 8 biți sau 4 biți.
Testează-teTest explicativ pentru modelele AI

Ce sa întâmplat

NVIDIA describes a set of JAX and Transformer Engine optimizations for dropless mixture-of-experts training, including grouped GEMM, fused expert-parallel dispatch and combine operations, NCCL EP, XLA multistream collectives, MXFP8 grouped , and host activation offloading. The company says an unoptimized DeepSeek-V3 training baseline reached 103 TFLOPS per GPU, compared with 1,068 TFLOPS per GPU after optimization, and that end-to-end throughput improved by about 10x. NVIDIA also reports 97% scaling efficiency at 1,024 GPUs. The optimized components are included in the NVIDIA NGC MaxText container, with instructions for testing and reproducing the configuration.

NVIDIA’s post focuses on dropless mixture-of-experts training, in which every token is processed by its selected expert even when routing creates uneven expert loads. Unlike capacity-based approaches that trim overflow or pad experts to fixed sizes, dropless training requires kernels and communication paths to handle variable token counts. NVIDIA says these irregular shapes produce ragged tensors and can leave GPUs waiting on dispatch, all-to-all communication, and expert computation.

The reported changes operate at several layers. Transformer Engine’s grouped GEMM processes variable-length expert groups in one kernel call, while fused dispatch and combine operations and the NCCL EP backend address movement of tokens between GPUs. NVIDIA also cites XLA multistream collectives, host activation offloading, and MXFP8 grouped as contributors to the result.

The company reports that its DeepSeek-V3 671B configuration increased a baseline from 103 to 1,068 TFLOPS per GPU and produced an approximately 10x end-to-end throughput gain. It also reports 97% efficiency at 1,024 GPUs. These figures are presented by NVIDIA as observations from its optimized stack, not as independently verified results.

The implementation is available to try through the NVIDIA NGC MaxText container with Transformer Engine built in. The post specifies ghcr.io/nvidia/jax:maxtext-2026-09-09 or newer and provides configuration guidance for DeepSeek-V3. It cautions that different models require different tuning. Pricing and licensing details are not provided.

Detalii sursa: developer.nvidia.com ↗

De ce contează

Large mixture-of-experts models reduce computation by activating only selected expert networks, but their uneven token routing creates difficult communication and memory bottlenecks. If NVIDIA’s results hold beyond its reported configuration, the changes could make training very large MoE systems more efficient without dropping routed tokens or padding every expert to a fixed capacity. That matters primarily to organizations operating large NVIDIA GPU clusters, where communication overhead and wasted computation can materially affect training time and cost. The results are NVIDIA’s own claims; the source does not provide independent benchmarking or a comparison with every alternative implementation.

MoE training is increasingly important for large AI models because conditional computation can reduce the active computation per token while retaining many parameters. The engineering difficulty shifts toward routing, irregular matrix operations, memory use, and interconnect traffic. Improving those parts could reduce the time and hardware required to train models at production scale.

The practical audience is specialized: researchers and companies already using JAX, MaxText, NVIDIA GPUs, and multi-GPU or multi-node infrastructure. The source does not establish that the same gains apply to smaller systems, non-NVIDIA hardware, dense models, or workloads outside the reported DeepSeek-V3 configuration.

NVIDIA’s figures should be treated as vendor-reported performance claims. The source supplies no independent tests, detailed cost analysis, or full methodology sufficient to determine how much of the improvement comes from each optimization.

Interactive Mechanism

Mecanism interactiv: cum funcționează de fapt

Explorați tehnologia care stau la baza acestei dezvoltări în mod interactiv.

Document Size:128K tokens
Needle Placement Depth (Location in document):50% into text
Attention Context Buffer Map:
Target Fact (50%)
Equivalent Pages~320Standard book pages
Retrieval Accuracy99.9%Needle recall score
RAM / KV Cache5.1 GBMemory overhead
Prompt CachingActive~80% discount on reuse
Core takeaway: Million-token context windows allow querying whole codebases or legal archives in one prompt. However, KV cache memory scales with context length, making prompt caching crucial for real-time production.
Verificare interactivă a conceptului+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Ce să urmărești în continuare

The main question is whether the reported gains reproduce across different MoE architectures, cluster layouts, model sizes, and GPU generations. NVIDIA plans to add NVFP4, fused with GEMM, and additional all-to-all overlap. Users should also watch whether these features become standard in JAX and MaxText workflows, and whether the optimizations require substantial model-specific tuning. The source documents a container-based access path but does not state pricing, licensing terms, support commitments, or general availability beyond the cited NGC container.

Reproduction across other MoE models and hardware configurations will indicate whether the reported gains are broadly transferable or tightly coupled to DeepSeek-V3 and NVIDIA’s cluster setup.

NVIDIA says future work will add NVFP4, fused with GEMM, and additional all-to-all overlap. Those additions could further affect the performance and memory tradeoffs of the stack.

The source recommends tracking step time, TFLOPS per GPU, model FLOP utilization, grouped GEMM latency, and dispatch/combine latency. Those measurements will help separate compute gains from communication improvements.

The post does not state the container’s price, licensing conditions, support level, or whether all cited components are equally mature for production use.

Ghiduri și chestionare conexe

Modelele AI explicateTransformatoareAntrenament AIViitorul IATestați ceea ce știți — încercați un test AI gratuitCăutați un termen AI în glosarul nostruUrmați instrumentul de urmărire a lansării modelului AI
Ai găsit asta util?