뉴스로 돌아가기
제품AI Understanding 브리핑

NVIDIA introduces NVLink Fusion for custom AI accelerators

NVIDIA says NVLink Fusion will let hyperscalers and AI companies connect custom XPUs to its networking, rack, cooling, power and software infrastructure, potentially reducing the complexity and time required to deploy semi-custom AI systems.

5 min readRead the primary source
Primary-source image accompanying NVIDIA introduces NVLink Fusion for custom AI accelerators
기본 소스 문서녹음된 소스
출판사
blogs.nvidia.com
소스 링크
blogs.nvidia.comhttps://blogs.nvidia.com/blog/nvlink-fusion-xpu-ai-factory/
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

훈련 후
명령어 튜닝, 선호도 최적화, 안전 튜닝 등 사전 학습 이후 적용되는 학습 단계입니다.
추론
훈련된 모델이 예측 또는 출력을 생성하는 런타임 단계입니다.
매개변수
출력에 영향을 미치는 모델 내부의 학습된 가중치입니다.
자신을 테스트해 보세요AI 모델 설명 퀴즈

무슨 일이 일어났나요?

NVIDIA says NVLink Fusion is a platform for connecting custom XPUs to its established AI-factory infrastructure. The program covers scale-up networking, rack architecture, cooling, power delivery, software, manufacturing and supplier integration.

NVIDIA’s August 24 blog presents NVLink Fusion as a way for hyperscalers and AI-native companies to use custom accelerators without designing an entire AI data-center platform from scratch. The company says AI infrastructure must be treated as a continuously operating factory whose economics depend on delivered output, energy use, cost per token, utilization and uptime. On that view, an XPU design is only one part of the deployment problem; networking, racks, cooling, power, software and suppliers also determine whether the system can operate at scale.

The central offering is access to NVIDIA’s NVLink scale-up domain. NVIDIA says sixth-generation NVLink can connect 72 XPUs with high-bandwidth, low-latency communication, and claims that XPU-to-XPU latency is three times lower and packet rate ten times higher than alternatives based on off-the-shelf Ethernet.

The company also describes NVLink-C2C, which connects XPUs to NVIDIA Vera CPUs or other ecosystem CPUs, as delivering up to six times the energy efficiency of a PCIe interface. These are NVIDIA’s stated figures; the source does not identify test conditions, independent measurements or a specific third-party XPU configuration.

NVLink Fusion also links custom silicon to NVIDIA’s MGX rack-scale architecture and associated supply chain. The blog says adopters can use established designs for compute and switch trays, cooling, power delivery and management, while manufacturing partners handle design and integration. NVIDIA says XPU- and GPU-based systems can share rack footprints, networking, cooling, power delivery and management systems, allowing operators to begin facility planning before deciding the final accelerator mix.

The platform is presented as including software as well as hardware. NVIDIA names NCCL for distributed workloads, Dynamo and NIXL for disaggregation, and Mission Control for cluster management, telemetry and debugging. The page includes supportive comments from executives at Intel, QCT and Quanta Computer, MediaTek, GUC and Annapurna Labs, but it does not identify a completed deployment using a customer’s XPU or provide a launch schedule for individual systems.

소스 세부정보: blogs.nvidia.com

왜 중요한가요?

Custom accelerators can be attractive for specialized AI workloads, but building the surrounding infrastructure is a major engineering and deployment challenge. NVIDIA’s approach aims to let customers differentiate their silicon while reusing much of the platform around it.

The practical significance is that custom AI silicon often requires custom infrastructure. According to NVIDIA, teams developing XPUs must integrate high-speed CPU and scale-up interfaces, validate networking, design compute and switch trays, engineer cooling and power, manage security and storage, and coordinate multiple suppliers. Reusing a mature platform could reduce the amount of work required before a custom accelerator reaches a data center.

This matters because AI systems increasingly combine different kinds of computation. The source specifically cites trillion- models, mixture-of-experts systems and agentic AI as workloads where communication between accelerators can affect utilization and cost per token. It also describes GPUs, XPUs, CPUs and LPUs being used together for training, , reasoning, retrieval and serving. If the infrastructure can accommodate several accelerator types, operators may have more flexibility to match hardware to workloads.

NVIDIA’s proposal could also shift where competition occurs. Customers could focus their engineering effort on a targeted XPU design while using NVIDIA’s networking, rack architecture and software stack for the rest of the system. That may lower deployment risk and accelerate access to manufacturing capacity, but it could also make custom-accelerator builders more dependent on NVIDIA’s interfaces, software and supply chain. The source does not quantify either the cost savings or the degree of technical dependence.

The facility-planning argument is another concrete consideration. NVIDIA says power procurement, cooling, rack layouts and network architecture are planned before final silicon is available, and that a facility designed around one chip can become a schedule risk. A common rack and infrastructure design could allow operators to defer some silicon decisions and reprovision capacity as demand or supply changes. That benefit remains a company claim rather than a demonstrated result in the material provided.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
대화형 개념 확인+10 Points
AI Models Explained Quiz

What is the best response when AI Models Explained makes a mistake in production?

다음에 무엇을 볼 것인가

The source does not provide pricing, customer deployment dates, independent performance testing or evidence from a production XPU system. The key question is whether third-party XPUs can deliver meaningful gains in real workloads while retaining the operational benefits NVIDIA describes.

The most important evidence would be a named customer deployment using a non-NVIDIA XPU with NVLink Fusion in production. The source describes the program and quotes ecosystem participants, but it does not establish that a third-party XPU is operating at scale, meeting uptime targets or achieving the claimed communication and energy characteristics.

Independent testing will be needed to assess the performance claims. NVIDIA gives comparative figures for latency, packet rate and CPU-interconnect energy efficiency, but it does not state the exact competing hardware, software versions, workload mix, measurement methodology or whether the results apply equally to custom XPUs. Results from NVIDIA’s own AI benchmarks are not the same as independent validation of NVLink Fusion with customer silicon.

Availability and commercial terms are also unknown. The blog does not state when NVLink Fusion hardware, interfaces or software support will be generally available, which XPU designs are eligible, what integration work customers must perform, or how licensing and supply arrangements are structured. It also does not name specific hyperscalers or AI-native companies that have committed to deployment.

Finally, the roadmap should be treated cautiously. NVIDIA mentions future NVLink configurations supporting domains of up to 1,152 accelerators and co-packaged optics, but the page does not give delivery dates or demonstrate those systems. Coverage should track whether the proposed common infrastructure produces measurable benefits in cost, power, utilization, serviceability and time to deployment outside NVIDIA’s own reference environments.

관련 가이드 및 퀴즈

AI 모델 설명AI 에이전트AI 트레이닝AI의 미래알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.
이것이 유용하다고 생각하시나요?