3 Sources
[1]
Nvidia custom 'NVHBM' promises 30% higher bandwidth, 15% lower power than commodity HBM4e -- custom base die and PHY will be available to NVLink Fusion partners
Nvidia's NVLink Fusion program gives the company's partners the building blocks necessary to connect custom chips with the NVLink scale-up domain used to join many separate processors into a single coherent system like the Vera Rubin NVL72 rack-scale accelerator. Today, Nvidia is adding a new
[2]
NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory
Amazon's Annapurna Labs will be the first to collaborate on NVHBM technology alongside NVLink Fusion. The next wave of AI is placing new demands on infrastructure. As AI agents and trillion-parameter workloads become mainstream, the performance of AI infrastructure depends not only on compute,
[3]
NVIDIA Develops Custom "NVHBM" Memory For AI, Claiming 30% More Bandwidth and 15% Lower Power Than HBM4E
NVIDIA has just announced its custom HBM solution called NVHBM, which will be used in future GPUs, offering higher bandwidth & lower power. NVHBM Is A Custom HBM Solution Designed By NVIDIA & Amazon's Annapurna Labs To Speed Up DRAM Capabilities For Future Agentic & Physical AI Workloads The
Share
Copy Link
Nvidia has introduced NVHBM, a custom high-bandwidth memory solution that promises 30% higher bandwidth and 15% lower power consumption than standard HBM4e. The technology moves the memory controller into the HBM base die, freeing up to 25% more die area for compute. Amazon's Annapurna Labs will be the first partner to adopt NVHBM for future AWS infrastructure designs.

Nvidia has announced NVHBM, a custom high-bandwidth memory solution designed to address the escalating demands of AI infrastructure as trillion-parameter workloads and AI agents become mainstream
1
2
. The technology represents an expansion of the NVLink Fusion program, which provides partners with building blocks to connect custom chips within Nvidia's rack-scale platform architecture2
. NVHBM delivers 30% higher bandwidth per stack compared to standard HBM4e, directly translating to higher tokens-per-second rates for AI inference workloads that are memory-bandwidth-bound1
3
.Traditional HBM architectures incorporate the memory controller into the primary silicon die on the package, consuming valuable real estate that could be dedicated to compute functions
2
. NVHBM fundamentally changes this approach by integrating Nvidia's custom memory controller directly into the base die of the 3D HBM stack1
3
. This architectural shift frees up to 25% more area on XPU compute dies, allowing designers to add up to 30% more compute capabilities on the primary silicon die1
2
. The technology also simplifies interposer routing used to join multiple chips together in designs using advanced packaging techniques1
.NVHBM achieves 15% lower power consumption compared to commodity HBM4e, addressing a critical concern for AI infrastructure operators
1
2
. These power savings can be reallocated into more functional units for custom AI chip designs, translated into higher sustained performance within the same power budget, or banked for improved performance-per-watt metrics1
. For hyperscalers deploying thousands of AI accelerators, these savings multiply significantly when moving massive data structures like model weights and KV caches1
. The energy saved on data movement can support larger numbers of accelerators within the same fixed power envelope1
.Related Stories
Amazon's Annapurna Labs will be the first partner to collaborate on NVHBM technology as part of its broader work with Nvidia around NVLink Fusion
2
3
. Nafea Bshara, vice president of Annapurna Labs at Amazon, stated that "NVHBM represents a new architectural approach to advancing high-bandwidth memory performance and efficiency" and expressed anticipation for how "this technology collaboration to benefit future AWS infrastructure designs"2
3
. Annapurna's next-generation Trainium 4 chips will support NVLink Fusion, allowing Amazon chips and Nvidia GPUs to work together within a common rack-scale architecture1
2
.Nvidia is establishing a standard NVHBM implementation that will be validated and offered by leading memory vendors, reducing the engineering effort required to integrate and qualify memory across multiple suppliers
2
3
. This standardization provides NVLink Fusion customers with a faster path for bringing custom AI chips to market compared to implementing commodity HBM from the ground up1
2
. The technology will be available exclusively to Nvidia's custom silicon partners through the NVLink Fusion program, which also provides access to NVLink chiplets, NVLink-C2C, NVLink Switches, and Nvidia MGX systems and racks2
. Nvidia has indicated that NVHBM will be incorporated into future GPUs, with speculation pointing to the Feynman GPU generation expected in 2028 as a potential first adopter3
.Summarized by
Navi
[1]
19 Mar 2025•Technology

11 Dec 2024•Technology

10 Feb 2026•Technology

1
Technology

2
Technology

3
Science and Research
