Ku laabo Warka
Hal-abuurnimoAI Understanding warbixin kooban

NVIDIA waxay soo tebisay 10x tabobar MoE ah oo degdeg ah oo aan dhibic lahayn gudaha JAX

NVIDIA waxay leedahay Engineerkeeda Transformer iyo hagaajinta JAX waxay kordhisay wax soo saarka tababarka DeepSeek-V3 671B qiyaastii 10x oo ay ku guulaysatay 97% waxtarka guud ahaan 1,024 GPUs. Jidka la hagaajiyay waxa lagu daray weelka NVIDIA NGC MaxText ee ku taariikhaysan Sebtembar 9, 2026, ama ka cusub.

4 min readRead the primary source
Source-provided image accompanying NVIDIA reports 10x faster dropless MoE training in JAX
Dukumeentiga isha aasaasiga ahIsha la duubay
Daabacaha
developer.nvidia.com
Xidhiidhka isha
developer.nvidia.comhttps://developer.nvidia.com/blog/accelerating-dropless-moe-training-in-jax-with-nvidia-transformer-engine/
Nooca isha
Dukumeentiga aasaasiga ah - ogeysiis rasmi ah, warqad, xereyn, ama bogga xisbiga koowaad waxaan si toos ah u akhrinay.
Dulucda sheekadaKu fahan tan 60 ilbiriqsi gudahood

Halkan ka bilow

Qodobbada muhiimka ah

Isku dhafka Khubarada (MoE)
Naqshad leh shabakado-hoosaadyo gaar ah oo khubaro la doortay oo keliya ay ku shaqeeyaan wixii talo-gelin ah.
Xusuusta (Xusuusta Wakiilka)
Macnaha guud ee la kaydiyay wakiilka AI wuxuu isticmaalaa dhammaan tillaabooyinka ama fadhiyada si uu u horumariyo sii wadida.
Tirada
Miisaanka moodeelka oo loo beddelo qaababka saxda ah ee hoose sida 8-bit ama 4-bit.
Is tijaabiMoodooyinka AI Kedis La Sharaxay

Maxaa dhacay

NVIDIA describes a set of JAX and Transformer Engine optimizations for dropless mixture-of-experts training, including grouped GEMM, fused expert-parallel dispatch and combine operations, NCCL EP, XLA multistream collectives, MXFP8 grouped , and host activation offloading. The company says an unoptimized DeepSeek-V3 training baseline reached 103 TFLOPS per GPU, compared with 1,068 TFLOPS per GPU after optimization, and that end-to-end throughput improved by about 10x. NVIDIA also reports 97% scaling efficiency at 1,024 GPUs. The optimized components are included in the NVIDIA NGC MaxText container, with instructions for testing and reproducing the configuration.

NVIDIA’s post focuses on dropless mixture-of-experts training, in which every token is processed by its selected expert even when routing creates uneven expert loads. Unlike capacity-based approaches that trim overflow or pad experts to fixed sizes, dropless training requires kernels and communication paths to handle variable token counts. NVIDIA says these irregular shapes produce ragged tensors and can leave GPUs waiting on dispatch, all-to-all communication, and expert computation.

The reported changes operate at several layers. Transformer Engine’s grouped GEMM processes variable-length expert groups in one kernel call, while fused dispatch and combine operations and the NCCL EP backend address movement of tokens between GPUs. NVIDIA also cites XLA multistream collectives, host activation offloading, and MXFP8 grouped as contributors to the result.

The company reports that its DeepSeek-V3 671B configuration increased a baseline from 103 to 1,068 TFLOPS per GPU and produced an approximately 10x end-to-end throughput gain. It also reports 97% efficiency at 1,024 GPUs. These figures are presented by NVIDIA as observations from its optimized stack, not as independently verified results.

The implementation is available to try through the NVIDIA NGC MaxText container with Transformer Engine built in. The post specifies ghcr.io/nvidia/jax:maxtext-2026-09-09 or newer and provides configuration guidance for DeepSeek-V3. It cautions that different models require different tuning. Pricing and licensing details are not provided.

Faahfaahinta isha: developer.nvidia.com ↗

Maxay muhiim u tahay

Large mixture-of-experts models reduce computation by activating only selected expert networks, but their uneven token routing creates difficult communication and memory bottlenecks. If NVIDIA’s results hold beyond its reported configuration, the changes could make training very large MoE systems more efficient without dropping routed tokens or padding every expert to a fixed capacity. That matters primarily to organizations operating large NVIDIA GPU clusters, where communication overhead and wasted computation can materially affect training time and cost. The results are NVIDIA’s own claims; the source does not provide independent benchmarking or a comparison with every alternative implementation.

MoE training is increasingly important for large AI models because conditional computation can reduce the active computation per token while retaining many parameters. The engineering difficulty shifts toward routing, irregular matrix operations, memory use, and interconnect traffic. Improving those parts could reduce the time and hardware required to train models at production scale.

The practical audience is specialized: researchers and companies already using JAX, MaxText, NVIDIA GPUs, and multi-GPU or multi-node infrastructure. The source does not establish that the same gains apply to smaller systems, non-NVIDIA hardware, dense models, or workloads outside the reported DeepSeek-V3 configuration.

NVIDIA’s figures should be treated as vendor-reported performance claims. The source supplies no independent tests, detailed cost analysis, or full methodology sufficient to determine how much of the improvement comes from each optimization.

Interactive Mechanism

Farsamaynta Is-dhexgalka: Sida Dhabta Ay U Shaqeyso

U baadh tignoolajiyada hoose ee ka dambeeya horumarkan si isdhexgal leh.

Document Size:128K tokens
Needle Placement Depth (Location in document):50% into text
Attention Context Buffer Map:
Target Fact (50%)
Equivalent Pages~320Standard book pages
Retrieval Accuracy99.9%Needle recall score
RAM / KV Cache5.1 GBMemory overhead
Prompt CachingActive~80% discount on reuse
Core takeaway: Million-token context windows allow querying whole codebases or legal archives in one prompt. However, KV cache memory scales with context length, making prompt caching crucial for real-time production.
Hubinta Fikradda Is-dhexgalka+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Maxaa la daawan doona xiga

The main question is whether the reported gains reproduce across different MoE architectures, cluster layouts, model sizes, and GPU generations. NVIDIA plans to add NVFP4, fused with GEMM, and additional all-to-all overlap. Users should also watch whether these features become standard in JAX and MaxText workflows, and whether the optimizations require substantial model-specific tuning. The source documents a container-based access path but does not state pricing, licensing terms, support commitments, or general availability beyond the cited NGC container.

Reproduction across other MoE models and hardware configurations will indicate whether the reported gains are broadly transferable or tightly coupled to DeepSeek-V3 and NVIDIA’s cluster setup.

NVIDIA says future work will add NVFP4, fused with GEMM, and additional all-to-all overlap. Those additions could further affect the performance and memory tradeoffs of the stack.

The source recommends tracking step time, TFLOPS per GPU, model FLOP utilization, grouped GEMM latency, and dispatch/combine latency. Those measurements will help separate compute gains from communication improvements.

The post does not state the container’s price, licensing conditions, support level, or whether all cited components are equally mature for production use.

Tilmaamaha la xidhiidha & su'aalaha

Moodooyinka AI ayaa la sharaxayTransformersTababarka AIMustaqbalka AITijaabi waxaad taqaan - isku day kedis AI oo bilaash ahKa raadi erey AI qaamuuskeenaRaac qaabka AI raadraaca sii deynta
Tan faa'iido ma u heshay?