d-Matrix adopts Nvidia NVLink Fusion to accelerate AI inference chip deployment

6 Sources

Share

AI chipmaker d-Matrix will integrate Nvidia's NVLink Fusion technology into its next-generation Raptor XPUs, enabling direct deployment within Nvidia's data-center systems. The Microsoft-backed startup aims to deliver ultralow-latency inference for chatbots and coding assistants by 2027.

d-Matrix Embraces Nvidia Infrastructure for Faster Market Entry

AI chipmaker d-Matrix announced it will integrate Nvidia's NVLink Fusion interconnect technology into its upcoming Raptor XPUs, joining a growing list of chipmakers licensing Nvidia chip-linking tech to accelerate deployment

1

2

. The move positions d-Matrix to tap into Nvidia's established AI infrastructure platform, including Nvidia MGX rack designs, Spectrum-X networking, and validated supply chains

3

. By adopting this approach, d-Matrix sidesteps the costly and time-consuming process of building separate rack architectures and scale-up networks from scratch.

"Demand for inference is soaring, but capital, time and energy remain finite," said Sid Sheth, cofounder and CEO of d-Matrix

4

. The partnership gives customers a lower-risk path to deploy specialized AI workloads at scale within liquid-cooled architectures already proven in data-center systems.

Raptor XPUs Target Memory Bandwidth Bottleneck

The Santa Clara, California-based startup specializes in AI inference chips designed to address memory bandwidth constraints that slow token generation in chatbots, coding assistants, and voice agents where speed matters most

5

. Each Raptor card features 32 GB of ultra-fast 3D-stacked DRAM delivering 100 TB/s of memory bandwidth—roughly 4.5 times the memory bandwidth of Nvidia's Rubin GPU

1

.

By bonding compute logic atop DRAM stacks, d-Matrix achieves in-memory compute performance that maintains memory bandwidth closer to SRAM while offering significantly higher capacity. This architecture allows d-Matrix to handle trillion-parameter models with far fewer chips compared to alternatives. Where similar workloads might require over 2,000 chips at 8-bit precision using competing technologies, a single d-Matrix system would need only about 64 chips (32 at 4-bit precision)

1

.

Rack-Scale XPU Deployment Planned for 2027

By the end of next year, d-Matrix expects to offer systems with up to 144 Raptor accelerators connected through a single all-to-all NVLink fabric

1

. These NVL144 racks will deliver approximately 2.3 TB of memory capacity—sufficient for models exceeding four trillion parameters at 4-bit precision—and about 7.2 petabytes per second of peak aggregate memory bandwidth. The Nvidia-compatible racks are expected to be available in 2027, with Raptor chips scheduled to complete their final design stage by the end of this year

2

.

Source: NVIDIA

Source: NVIDIA

d-Matrix plans to integrate Nvidia Vera CPUs, ConnectX-9 SuperNICs, BlueField-4 DPUs, and Spectrum-X Ethernet networking alongside NVLink Fusion

4

. The startup is also partnering with connectivity firm Astera Labs to build custom solutions ensuring fast data flow across the system

5

.

Nvidia's Expanding Ecosystem Strategy

NVLink Fusion extends Nvidia's platform openness to XPUs and CPUs, allowing silicon companies to focus on processor innovations while leveraging Nvidia infrastructure to deploy them at AI factory scale

3

. The common rack architecture enables data centers to support GPUs, CPUs, and XPUs without requiring separate rack designs for each processor type.

d-Matrix joins MediaTek, Marvell, Qualcomm, Arm, Fujitsu, and Amazon Web Services in adopting Nvidia's NVLink Fusion interconnects

1

. Nvidia has invested billions in incentives to expand this ecosystem—$3.5 billion in MediaTek and $2 billion in Marvell—though financial terms of the d-Matrix collaboration were not disclosed.

What This Means for AI Inference Markets

Microsoft has backed d-Matrix since its $110 million financing round in 2023. The startup, which shipped its first AI chip in November 2024, was valued at $2 billion when it raised $450 million last year

5

. As AI workloads shift from training to inference—running models for everyday use—specialized chips like Raptor XPUs offer potential advantages in low-latency AI services where response time drives user experience.

The partnership reflects broader industry trends toward disaggregated inference architectures. d-Matrix chips can be deployed standalone or alongside GPUs in heterogeneous configurations, using GPUs for compute-intensive prompt processing and Raptor accelerators for memory bandwidth-bound token generation

1

. This flexibility may appeal to enterprises seeking ultralow-latency inference capabilities at lower entry costs compared to alternatives requiring thousands of chips per deployment.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved