뉴스로 돌아가기
혁신AI Understanding 브리핑

NVIDIA는 JAX에서 10배 더 빠른 드롭리스 MoE 교육을 보고합니다.

NVIDIA는 Transformer Engine 및 JAX 최적화로 DeepSeek-V3 671B 교육 처리량이 약 10배 증가했으며 1,024개의 GPU에서 97%의 효율성을 달성했다고 밝혔습니다. 최적화된 경로는 2026년 9월 9일 이후 날짜의 NVIDIA NGC MaxText 컨테이너에 포함되어 있습니다.

4 min readRead the primary source
Source-provided image accompanying NVIDIA reports 10x faster dropless MoE training in JAX
기본 소스 문서녹음된 소스
출판사
developer.nvidia.com
소스 링크
developer.nvidia.comhttps://developer.nvidia.com/blog/accelerating-dropless-moe-training-in-jax-with-nvidia-transformer-engine/
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

전문가 혼합(MoE)
입력당 선택된 전문가만 실행하는 특수 하위 네트워크가 있는 아키텍처입니다.
메모리(에이전트 메모리)
AI 에이전트는 연속성을 향상하기 위해 여러 단계 또는 세션에서 사용하는 저장된 컨텍스트입니다.
양자화
모델 가중치를 8비트 또는 4비트와 같은 낮은 정밀도 형식으로 변환합니다.
자신을 테스트해 보세요AI 모델 설명 퀴즈

무슨 일이 일어났나요?

NVIDIA describes a set of JAX and Transformer Engine optimizations for dropless mixture-of-experts training, including grouped GEMM, fused expert-parallel dispatch and combine operations, NCCL EP, XLA multistream collectives, MXFP8 grouped , and host activation offloading. The company says an unoptimized DeepSeek-V3 training baseline reached 103 TFLOPS per GPU, compared with 1,068 TFLOPS per GPU after optimization, and that end-to-end throughput improved by about 10x. NVIDIA also reports 97% scaling efficiency at 1,024 GPUs. The optimized components are included in the NVIDIA NGC MaxText container, with instructions for testing and reproducing the configuration.

NVIDIA’s post focuses on dropless mixture-of-experts training, in which every token is processed by its selected expert even when routing creates uneven expert loads. Unlike capacity-based approaches that trim overflow or pad experts to fixed sizes, dropless training requires kernels and communication paths to handle variable token counts. NVIDIA says these irregular shapes produce ragged tensors and can leave GPUs waiting on dispatch, all-to-all communication, and expert computation.

The reported changes operate at several layers. Transformer Engine’s grouped GEMM processes variable-length expert groups in one kernel call, while fused dispatch and combine operations and the NCCL EP backend address movement of tokens between GPUs. NVIDIA also cites XLA multistream collectives, host activation offloading, and MXFP8 grouped as contributors to the result.

The company reports that its DeepSeek-V3 671B configuration increased a baseline from 103 to 1,068 TFLOPS per GPU and produced an approximately 10x end-to-end throughput gain. It also reports 97% efficiency at 1,024 GPUs. These figures are presented by NVIDIA as observations from its optimized stack, not as independently verified results.

The implementation is available to try through the NVIDIA NGC MaxText container with Transformer Engine built in. The post specifies ghcr.io/nvidia/jax:maxtext-2026-09-09 or newer and provides configuration guidance for DeepSeek-V3. It cautions that different models require different tuning. Pricing and licensing details are not provided.

소스 세부정보: developer.nvidia.com ↗

왜 중요한가요?

Large mixture-of-experts models reduce computation by activating only selected expert networks, but their uneven token routing creates difficult communication and memory bottlenecks. If NVIDIA’s results hold beyond its reported configuration, the changes could make training very large MoE systems more efficient without dropping routed tokens or padding every expert to a fixed capacity. That matters primarily to organizations operating large NVIDIA GPU clusters, where communication overhead and wasted computation can materially affect training time and cost. The results are NVIDIA’s own claims; the source does not provide independent benchmarking or a comparison with every alternative implementation.

MoE training is increasingly important for large AI models because conditional computation can reduce the active computation per token while retaining many parameters. The engineering difficulty shifts toward routing, irregular matrix operations, memory use, and interconnect traffic. Improving those parts could reduce the time and hardware required to train models at production scale.

The practical audience is specialized: researchers and companies already using JAX, MaxText, NVIDIA GPUs, and multi-GPU or multi-node infrastructure. The source does not establish that the same gains apply to smaller systems, non-NVIDIA hardware, dense models, or workloads outside the reported DeepSeek-V3 configuration.

NVIDIA’s figures should be treated as vendor-reported performance claims. The source supplies no independent tests, detailed cost analysis, or full methodology sufficient to determine how much of the improvement comes from each optimization.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Document Size:128K tokens
Needle Placement Depth (Location in document):50% into text
Attention Context Buffer Map:
Target Fact (50%)
Equivalent Pages~320Standard book pages
Retrieval Accuracy99.9%Needle recall score
RAM / KV Cache5.1 GBMemory overhead
Prompt CachingActive~80% discount on reuse
Core takeaway: Million-token context windows allow querying whole codebases or legal archives in one prompt. However, KV cache memory scales with context length, making prompt caching crucial for real-time production.
대화형 개념 확인+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

다음에 무엇을 볼 것인가

The main question is whether the reported gains reproduce across different MoE architectures, cluster layouts, model sizes, and GPU generations. NVIDIA plans to add NVFP4, fused with GEMM, and additional all-to-all overlap. Users should also watch whether these features become standard in JAX and MaxText workflows, and whether the optimizations require substantial model-specific tuning. The source documents a container-based access path but does not state pricing, licensing terms, support commitments, or general availability beyond the cited NGC container.

Reproduction across other MoE models and hardware configurations will indicate whether the reported gains are broadly transferable or tightly coupled to DeepSeek-V3 and NVIDIA’s cluster setup.

NVIDIA says future work will add NVFP4, fused with GEMM, and additional all-to-all overlap. Those additions could further affect the performance and memory tradeoffs of the stack.

The source recommends tracking step time, TFLOPS per GPU, model FLOP utilization, grouped GEMM latency, and dispatch/combine latency. Those measurements will help separate compute gains from communication improvements.

The post does not state the container’s price, licensing conditions, support level, or whether all cited components are equally mature for production use.

관련 가이드 및 퀴즈

AI 모델 설명트랜스포머AI 트레이닝AI의 미래알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.
이것이 유용하다고 생각하시나요?