Voltar às notícias
ProdutoInstruções AI Understanding

NVIDIA details NVHBM memory design for custom AI accelerators

NVIDIA says its NVHBM technology can provide up to 30% more memory bandwidth, reduce memory-interface area by up to 67%, and use 15% less HBM power than standard HBM4e. The claims extend NVIDIA’s NVLink Fusion effort to connect custom accelerators to its rack-scale AI infrastructure.

Por 5 min read
NVIDIA’s source figure comparing NVHBM with standard HBM4e, focused on memory-interface area and the silicon space available for AI compute.
A versão curta

NVIDIA says its NVHBM technology can provide up to 30% more memory bandwidth, reduce memory-interface area by up to 67%, and use 15% less HBM power than standard HBM4e. The claims extend NVIDIA’s NVLink Fusion effort to connect custom accelerators to its rack-scale AI infrastructure.

O que aconteceu

NVIDIA’s technical blog provides additional details about NVHBM, a custom HBM base-die technology designed for custom AI accelerators used with NVLink Fusion. The company says the design targets memory bandwidth, chip area, power consumption, and integration constraints in large-scale AI systems.

NVIDIA says NVHBM is a custom HBM base-die technology designed and validated with leading memory vendors. The company presents it as a component of NVLink Fusion, its platform for allowing hyperscalers and AI-native companies to connect custom XPUs and CPUs to NVIDIA’s scale-up and scale-out networking, software ecosystem, and MGX rack architecture. The stated goal is to reduce the integration and deployment work involved in building semi-custom AI infrastructure. The description therefore covers both the memory component and the broader platform context in which NVIDIA expects it to be used.

The central claim is higher memory bandwidth. NVIDIA says NVHBM provides up to 30% more bandwidth per memory stack than standard HBM4e. The company argues that this could help keep accelerator compute engines supplied with model weights, key-value-cache data, and activations during training and large-scale inference. It specifically connects the technology to memory-bound workloads and to distributed techniques such as expert parallelism, where data and model experts are spread across accelerators and must be synchronized across a rack. The claimed benefit is consequently presented as a way to address data movement during both local processing and distributed operation.

NVIDIA also describes physical-design changes. It says NVHBM moves the memory controller into the three-dimensional HBM stack and uses a custom physical interface, or PHY. Compared with the JEDEC HBM4e standard, the company claims up to 67% less PHY and support area, up to 80% more usable silicon across the layout, and as much as 30% more main-die silicon available for compute or other features. The source further claims up to 15% lower HBM power use and says the combined NVLink Fusion and NVHBM design could deliver a 30% overall end-to-end performance increase per XPU. These are NVIDIA’s stated figures; the article does not give test conditions or independent verification. The figures describe potential design advantages, but the source leaves their practical conditions and limits unspecified.

Leia a fonte primária: developer.nvidia.com

Por que isso importa

If the stated improvements hold in production workloads, custom accelerator designers could fit more compute into a fixed package, feed memory-bound AI applications more effectively, and reduce power pressure in large data centers. The source is a vendor technical-marketing document, however, and does not provide independent testing, pricing, availability, or named production deployments for NVHBM.

Memory access is a major design constraint for modern AI accelerators because computation can be limited by how quickly the chip receives weights, activations, and cached intermediate data. NVIDIA’s proposal addresses that constraint at the package level while NVLink Fusion addresses communication among accelerators. In principle, the combination could reduce two forms of data bottleneck: local memory delivery within an accelerator and synchronization across a rack-scale compute domain. That framing makes the proposal relevant to system design as well as to the memory interface itself.

The area claim matters because accelerator designers must divide limited silicon and package space among matrix engines, vector units, SRAM, caches, control logic, memory interfaces, and networking. NVIDIA says a narrower, more efficient memory interface could allow designers to devote more of that fixed area to workload-specific compute or other capabilities. That flexibility could be relevant to inference serving, recommendation systems, multimodal applications, and internal training systems, although the source does not show a completed chip using those options. The significance of the claim depends on how designers would use the recovered area within an actual accelerator package.

Power savings could have infrastructure consequences beyond a single accelerator. NVIDIA says 15% lower HBM power could create room for additional compute, higher sustained utilization, or less demand on cooling and power-delivery systems. It gives a specific scenario in which the savings could provide headroom for up to 15,000 additional XPUs in a one-gigawatt data center using 2,000-watt XPUs. That figure is a projection based on the company’s assumptions, not evidence that a data center has achieved those results. The source also does not quantify possible costs from manufacturing changes, software integration, supply constraints, or thermal behavior in real deployments. Those unanswered questions are important when translating a component-level claim into a facility-level outcome.

O que assistir a seguir

The important next evidence will be independent benchmarks, concrete accelerator designs, memory-vendor and customer disclosures, production timelines, and measured results on training and inference workloads. Readers should also watch whether NVLink Fusion creates practical interoperability benefits for heterogeneous racks or primarily strengthens NVIDIA’s infrastructure ecosystem.

Independent performance testing should establish whether the claimed bandwidth and end-to-end gains appear across representative training, inference, and agentic workloads rather than only selected memory-bound cases. Useful disclosures would include accelerator configurations, batch sizes, model types, network traffic, power measurements, baselines, and whether the comparison uses identical software and thermal conditions. Such information would help separate improvements attributable to the memory design from improvements attributable to the surrounding system.

Availability and adoption will determine whether NVHBM is a practical platform or primarily a design option. NVIDIA does not name the memory vendors involved, identify a shipping custom XPU using the technology, provide customer commitments, or state when NVHBM-based products will enter production. Follow-up announcements should clarify qualification status, manufacturing scale, packaging requirements, commercial terms, and how broadly the technology can be sourced. These details will indicate whether the design can move from a technical announcement into repeatable product availability.

The broader strategic question is how much control custom accelerator developers retain while adopting NVLink Fusion. NVIDIA says the platform can connect custom XPUs and CPUs with its networking, software, and rack architecture, and can integrate custom XPUs with GPUs for heterogeneous computing. That may reduce development complexity, but it may also deepen dependence on NVIDIA’s interconnect and infrastructure stack. Evidence from multiple customers, interoperable products, and real rack-level deployments will show whether the claimed flexibility is broadly available or mainly benefits partners inside NVIDIA’s ecosystem. The answer will shape how the technology is viewed by developers weighing customization against platform integration.

Guias e questionários relacionados

Modelos de IA explicadosTreinamento de IAFuturo da IATeste o que você sabe – experimente um teste gratuito de IAProcure um termo de IA em nosso glossário
Achou isso útil?