Back to News
ProductAI Understanding briefing

NVIDIA introduces NVLink Fusion for custom AI accelerators

NVIDIA says NVLink Fusion will let hyperscalers and AI companies connect custom XPUs to its networking, rack, cooling, power and software infrastructure, potentially reducing the complexity and time required to deploy semi-custom AI systems.

By 5 min read
Primary-source image accompanying NVIDIA introduces NVLink Fusion for custom AI accelerators
The short version

NVIDIA says NVLink Fusion will let hyperscalers and AI companies connect custom XPUs to its networking, rack, cooling, power and software infrastructure, potentially reducing the complexity and time required to deploy semi-custom AI systems.

What happened

NVIDIA says NVLink Fusion is a platform for connecting custom XPUs to its established AI-factory infrastructure. The program covers scale-up networking, rack architecture, cooling, power delivery, software, manufacturing and supplier integration.

NVIDIA’s August 24 blog presents NVLink Fusion as a way for hyperscalers and AI-native companies to use custom accelerators without designing an entire AI data-center platform from scratch. The company says AI infrastructure must be treated as a continuously operating factory whose economics depend on delivered output, energy use, cost per token, utilization and uptime. On that view, an XPU design is only one part of the deployment problem; networking, racks, cooling, power, software and suppliers also determine whether the system can operate at scale.

The central offering is access to NVIDIA’s NVLink scale-up domain. NVIDIA says sixth-generation NVLink can connect 72 XPUs with high-bandwidth, low-latency communication, and claims that XPU-to-XPU latency is three times lower and packet rate ten times higher than alternatives based on off-the-shelf Ethernet.

The company also describes NVLink-C2C, which connects XPUs to NVIDIA Vera CPUs or other ecosystem CPUs, as delivering up to six times the energy efficiency of a PCIe interface. These are NVIDIA’s stated figures; the source does not identify test conditions, independent measurements or a specific third-party XPU configuration.

NVLink Fusion also links custom silicon to NVIDIA’s MGX rack-scale architecture and associated supply chain. The blog says adopters can use established designs for compute and switch trays, cooling, power delivery and management, while manufacturing partners handle design and integration. NVIDIA says XPU- and GPU-based systems can share rack footprints, networking, cooling, power delivery and management systems, allowing operators to begin facility planning before deciding the final accelerator mix.

The platform is presented as including software as well as hardware. NVIDIA names NCCL for distributed workloads, Dynamo and NIXL for disaggregation, and Mission Control for cluster management, telemetry and debugging. The page includes supportive comments from executives at Intel, QCT and Quanta Computer, MediaTek, GUC and Annapurna Labs, but it does not identify a completed deployment using a customer’s XPU or provide a launch schedule for individual systems.

Read the primary source: blogs.nvidia.com

Why it matters

Custom accelerators can be attractive for specialized AI workloads, but building the surrounding infrastructure is a major engineering and deployment challenge. NVIDIA’s approach aims to let customers differentiate their silicon while reusing much of the platform around it.

The practical significance is that custom AI silicon often requires custom infrastructure. According to NVIDIA, teams developing XPUs must integrate high-speed CPU and scale-up interfaces, validate networking, design compute and switch trays, engineer cooling and power, manage security and storage, and coordinate multiple suppliers. Reusing a mature platform could reduce the amount of work required before a custom accelerator reaches a data center.

This matters because AI systems increasingly combine different kinds of computation. The source specifically cites trillion-parameter models, mixture-of-experts systems and agentic AI as workloads where communication between accelerators can affect utilization and cost per token. It also describes GPUs, XPUs, CPUs and LPUs being used together for training, post-training, reasoning, retrieval and serving. If the infrastructure can accommodate several accelerator types, operators may have more flexibility to match hardware to workloads.

NVIDIA’s proposal could also shift where competition occurs. Customers could focus their engineering effort on a targeted XPU design while using NVIDIA’s networking, rack architecture and software stack for the rest of the system. That may lower deployment risk and accelerate access to manufacturing capacity, but it could also make custom-accelerator builders more dependent on NVIDIA’s interfaces, software and supply chain. The source does not quantify either the cost savings or the degree of technical dependence.

The facility-planning argument is another concrete consideration. NVIDIA says power procurement, cooling, rack layouts and network architecture are planned before final silicon is available, and that a facility designed around one chip can become a schedule risk. A common rack and infrastructure design could allow operators to defer some silicon decisions and reprovision capacity as demand or supply changes. That benefit remains a company claim rather than a demonstrated result in the material provided.

What to watch next

The source does not provide pricing, customer deployment dates, independent performance testing or evidence from a production XPU system. The key question is whether third-party XPUs can deliver meaningful gains in real workloads while retaining the operational benefits NVIDIA describes.

The most important evidence would be a named customer deployment using a non-NVIDIA XPU with NVLink Fusion in production. The source describes the program and quotes ecosystem participants, but it does not establish that a third-party XPU is operating at scale, meeting uptime targets or achieving the claimed communication and energy characteristics.

Independent testing will be needed to assess the performance claims. NVIDIA gives comparative figures for latency, packet rate and CPU-interconnect energy efficiency, but it does not state the exact competing hardware, software versions, workload mix, measurement methodology or whether the results apply equally to custom XPUs. Results from NVIDIA’s own AI inference benchmarks are not the same as independent validation of NVLink Fusion with customer silicon.

Availability and commercial terms are also unknown. The blog does not state when NVLink Fusion hardware, interfaces or software support will be generally available, which XPU designs are eligible, what integration work customers must perform, or how licensing and supply arrangements are structured. It also does not name specific hyperscalers or AI-native companies that have committed to deployment.

Finally, the roadmap should be treated cautiously. NVIDIA mentions future NVLink configurations supporting domains of up to 1,152 accelerators and co-packaged optics, but the page does not give delivery dates or demonstrate those systems. Coverage should track whether the proposed common infrastructure produces measurable benefits in cost, power, utilization, serviceability and time to deployment outside NVIDIA’s own reference environments.

Related guides & quizzes

AI Models ExplainedAI AgentsAI TrainingFuture of AITest what you know — try a free AI quizLook up an AI term in our glossary
Found this useful?