뉴스로 돌아가기
제품AI Understanding 브리핑

eTeknix는 Nvidia가 AI 가속기용 NVHBM 메모리 아키텍처를 도입했다고 보고합니다.

eTeknix는 Nvidia가 대규모 AI 시스템의 대역폭과 전력 효율성을 향상시키기 위한 고대역폭 메모리 설계인 NVHBM을 도입했다고 보고했습니다. 아웃렛은 Amazon가 미래 Trainium4 인프라에서 이를 사용할 계획이라고 밝혔고, Nvidia는 Feynman GPU가 2028년에 이 디자인을 채택할 것으로 예상하고 있습니다.

6 min readRead the linked source
Source-provided image accompanying eTeknix reports Nvidia introduced NVHBM memory architecture for AI accelerators
소스 참조녹음된 소스
출판사
eteknix.com
소스 링크
eteknix.comhttps://www.eteknix.com/nvidia-introduces-its-own-nvhbm-memory-for-artificial-intelligence/
소스 유형
연결된 소스 — 기본 소스 상태가 설정되지 않았습니다.
또한 인용됨

마지막으로 수정된 스토리

맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

메모리(에이전트 메모리)
AI 에이전트는 연속성을 향상하기 위해 여러 단계 또는 세션에서 사용하는 저장된 컨텍스트입니다.
벤치마크
모델 성능을 측정하고 비교하는 데 사용되는 표준화된 테스트 또는 데이터 세트입니다.
추론
훈련된 모델이 예측 또는 출력을 생성하는 런타임 단계입니다.
자신을 테스트해 보세요AI 모델 설명 퀴즈

출간 이후 달라진 점

  1. 처음 출판됨
  2. This technical NVIDIA Developer update materially advances the same NVLink Fusion and custom NVHBM event represented by the canonical update. It adds detailed company claims about NVHBM’s architecture and potential benefits: up to 30% more memory bandwidth than standard HBM4e, up to 67% lower PHY and support area, up to 30% more main-die silicon for compute or other features, up to 15% lower HBM power use, and a claimed 30% end-to-end performance increase per XPU when combined with NVLink Fusion.
  3. This is the same NVHBM product event covered by the canonical Nvidia NVLink Fusion update. eTeknix adds reported performance claims of more than 30% higher bandwidth, 15% lower power use, up to 25% processor-silicon savings and up to 80% package-silicon recovery, along with reported future adoption plans involving Amazon Trainium4 and Nvidia’s 2028 Feynman architecture. Those details are attributed to eTeknix and are not independently confirmed in the supplied material.

무슨 일이 일어났나요?

eTeknix reports that Nvidia introduced NVHBM, a high-bandwidth memory architecture designed with Amazon’s Annapurna Labs. The outlet says the design moves the memory controller from the main accelerator into the memory package’s base die and is intended for AI systems with very large models.

eTeknix reports that Nvidia has introduced NVHBM, its own high-bandwidth memory architecture for artificial-intelligence workloads. The article says Nvidia developed the design together with Annapurna Labs, Amazon’s chip division. The reported target is a class of AI systems whose models contain trillions of parameters, where moving data quickly between processors and memory can limit performance and increase power consumption. The source presents NVHBM as an infrastructure technology for AI accelerators rather than as a general-purpose consumer memory product. It does not provide a product price, shipping date or complete technical specification.

According to eTeknix, the distinguishing design choice is to place the memory controller directly in the base die of the high-bandwidth memory rather than keeping it inside the main processor, which the article calls an XPU. The report acknowledges that putting the controller in the base die is not unique to Nvidia and says Samsung, SK Hynix and Micron are exploring similar approaches. Nvidia’s claimed implementation, however, is presented as a way to reduce the amount of controller logic required on the accelerator and to narrow the physical memory interface. The source says this could simplify the interposer, the package component that connects the processor and memory.

The outlet reports several figures attributed to Nvidia: more than 30% additional bandwidth and 15% lower power use compared with HBM4E. It also says moving the controller can free up as much as 25% of the main chip’s silicon area, while reducing the interface width can recover as much as 80% of usable silicon in the package. These figures are company claims as reported by eTeknix; the supplied material contains no independent , test methodology, product sample or third-party confirmation. eTeknix further reports that Nvidia wants to standardize the specification across multiple memory makers so partners in its NVLink Fusion ecosystem can test and integrate custom accelerators more easily.

eTeknix says Amazon is among the first companies planning to adopt NVHBM. The article reports that Amazon intends to combine the memory architecture with Nvidia’s NVLink interconnect in future AWS Trainium4 processors, creating a rack-level architecture that brings Amazon chips and Nvidia GPUs together. On Nvidia’s side, the report says NVHBM is expected to debut with the Feynman GPU architecture, which is planned for 2028. Those are future plans as described by the outlet. The report does not establish present-day availability, completed hardware, customer deployments or a firm commercial launch schedule.

소스 세부정보: eteknix.com ↗

왜 중요한가요?

The reported design targets a central infrastructure constraint in AI: moving data between memory and processors efficiently. If the reported gains are achieved in production, NVHBM could give accelerator designers more bandwidth, lower power use and additional silicon area for computing resources.

Memory has become a limiting part of AI accelerator design because processors can only work efficiently when the data they need arrives quickly enough. eTeknix’s report frames NVHBM as an attempt to address that problem at the package level. The claimed combination of higher bandwidth and lower power use would matter most in large-scale training and systems, where many accelerators operate together and memory movement can consume substantial resources. The source does not quantify how these claimed improvements would translate into faster model training, lower operating costs or better end-user services.

The reported silicon savings could also affect how future accelerators are designed. If up to 25% of processor area is no longer needed for the memory controller, designers could potentially devote that space to additional computing resources, although eTeknix does not report a specific chip that does so. Its claim that up to 80% of usable package silicon can be recovered likewise describes design headroom, not a demonstrated performance result. The practical value will depend on how much of the recovered area can actually be used, whether the package remains affordable and whether the memory can be produced at the required scale.

The standardization plan matters because advanced memory is useful only when processors, packaging, firmware and manufacturing processes work together. eTeknix reports that Nvidia wants multiple memory makers to support a common NVHBM specification, which could reduce integration work for companies building custom accelerators through NVLink Fusion. The proposed Amazon connection would extend the significance beyond Nvidia’s own GPUs if Trainium4 eventually combines with Nvidia networking or accelerator infrastructure. Still, the supplied report does not document agreements with named memory suppliers, completed interoperability testing or commitments from cloud customers beyond the future Amazon adoption it describes.

The story therefore concerns more than a new memory label. It describes Nvidia trying to shape an important layer of the AI hardware stack while connecting that layer to its interconnect ecosystem and to Amazon’s processor roadmap. That could influence how custom AI systems are built and how much control Nvidia retains over surrounding components. The public impact remains prospective: the article supplies no evidence that NVHBM has yet changed available cloud capacity, accelerator pricing, energy consumption or access to AI services.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Document Size:128K tokens
Needle Placement Depth (Location in document):50% into text
Attention Context Buffer Map:
Target Fact (50%)
Equivalent Pages~320Standard book pages
Retrieval Accuracy99.9%Needle recall score
RAM / KV Cache5.1 GBMemory overhead
Prompt CachingActive~80% discount on reuse
Core takeaway: Million-token context windows allow querying whole codebases or legal archives in one prompt. However, KV cache memory scales with context length, making prompt caching crucial for real-time production.
대화형 개념 확인+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

다음에 무엇을 볼 것인가

The main unknowns are commercial availability, manufacturing partners, independent performance testing and the timing of deployments. eTeknix says Amazon plans future Trainium4 integration and Nvidia expects a 2028 Feynman debut, but the supplied report does not establish that either product is shipping or that the claimed benefits have been independently verified.

The first issue to watch is whether Nvidia publishes a full NVHBM specification and whether independent hardware analysts or customers test the reported bandwidth and power figures. The supplied report gives comparisons with HBM4E but does not describe workloads, memory capacities, operating conditions or measurement methods. Without those details, the percentages cannot be used to predict real-world gains for a particular model or data center. Independent testing would also clarify whether the improvements apply broadly or only to carefully selected configurations.

The second issue is supply and interoperability. Nvidia’s stated plan, as reported by eTeknix, is to standardize NVHBM across several memory makers, but the article does not name participating suppliers or identify manufacturing timelines. Future adoption will depend on whether those companies produce compatible parts in volume and whether packaging capacity can keep pace with AI accelerator demand. It will also be important to see whether custom-accelerator partners can use the design without becoming more dependent on Nvidia-specific interfaces or certification processes.

The third issue is the timing of the reported product roadmaps. eTeknix says Amazon plans to integrate NVHBM and NVLink into future Trainium4 processors and that Nvidia expects the technology to appear with Feynman GPUs in 2028. The report does not say when Trainium4 will launch, whether the two companies’ products will be available together, or whether the 2028 expectation is a firm commitment. Those dates should therefore be treated as forward-looking plans rather than current availability.

Finally, readers should watch for evidence of actual deployment effects: lower power use at system level, higher throughput for defined AI workloads, changes in accelerator density or reduced package complexity. None of those outcomes is established by this article. The meaningful unknowns include cost, yield, thermal behavior, memory capacity, compatibility with existing systems and the extent to which other suppliers’ competing base-die designs reach market. Until those questions are answered, NVHBM is a consequential reported hardware direction, but its practical advantage remains unconfirmed.

관련 가이드 및 퀴즈

AI 모델 설명AI 트레이닝트랜스포머AI의 미래알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.

업데이트 및 수정

이 정식 스토리는 진행 중인 이벤트가 실질적으로 변경될 때 업데이트됩니다. URL과 원래 출판 날짜는 절대 변경되지 않습니다.

  • This is the same NVHBM product event covered by the canonical Nvidia NVLink Fusion update. eTeknix adds reported performance claims of more than 30% higher bandwidth, 15% lower power use, up to 25% processor-silicon savings and up to 80% package-silicon recovery, along with reported future adoption plans involving Amazon Trainium4 and Nvidia’s 2028 Feynman architecture. Those details are attributed to eTeknix and are not independently confirmed in the supplied material.
  • This technical NVIDIA Developer update materially advances the same NVLink Fusion and custom NVHBM event represented by the canonical update. It adds detailed company claims about NVHBM’s architecture and potential benefits: up to 30% more memory bandwidth than standard HBM4e, up to 67% lower PHY and support area, up to 30% more main-die silicon for compute or other features, up to 15% lower HBM power use, and a claimed 30% end-to-end performance increase per XPU when combined with NVLink Fusion.
공개 수정 로그 보기
이것이 유용하다고 생각하시나요?