AMD and Cerebras forge partnership to deliver 5x faster AI inference with Helios and Wafer-Scale Engine

7 Sources

Share

AMD and Cerebras Systems announced a partnership to develop a disaggregated AI inference platform combining AMD's EPYC processors in Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine solutions. The collaboration promises up to 5x higher tokens per second per watt, positioning AMD to compete directly with Nvidia's $20 billion Groq acquisition while addressing the growing demand for ultra-low latency in AI applications.

AMD and Cerebras Partnership Targets Ultra-Low Latency AI

AMD and Cerebras Systems unveiled plans to develop a disaggregated AI inference platform that combines AMD's EPYC processors in Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine solutions

1

. The partnership, announced during AMD CEO Lisa Su's Advancing AI keynote Thursday, directly addresses a critical gap in AMD's portfolio and positions the company to compete with Nvidia's $20 billion acquisition of Groq in December

2

. The collaboration focuses on delivering ultra-low latency AI performance specifically designed for agentic workloads, where speed and responsiveness are paramount.

Source: Axios

Source: Axios

Workload Disaggregation in AI Drives Performance Gains

The new disaggregated AI inference platform assigns different portions of AI inference workloads to architectures optimized for them, promising up to 5x higher tokens per second per watt

1

. AMD Helios systems with Instinct GPUs handle the compute-heavy prompt processing stage and large context windows, while the Cerebras Wafer-Scale Engine takes over the memory-bandwidth-intensive token generation stage

1

. This approach inverts Nvidia's disaggregation strategy, where the cancelled Rubin CPX GPU was optimized for context/prefill while HBM-equipped GPUs handled generation

1

.

Cerebras Wafer-Scale Engine Leverages SRAM Memory Advantage

Cerebras CEO and cofounder Andrew Feldman emphasized the technological advantage of combining AMD's leadership in performance and memory capacity with Cerebras' dominance in SRAM and memory bandwidth

2

. Unlike GPUs that rely on HBM4, Cerebras' Wafer-Scale Engine uses on-chip SRAM that's orders of magnitude faster, enabling the company to achieve output speeds often exceeding 2,000 tokens per second

2

. The Wafer-Scale Engine places an entire AI supercomputer's worth of memory and compute onto a single, giant, interconnected sheet of silicon, where hundreds of thousands of compute cores and tens of gigabytes of SRAM connect seamlessly

5

. This architecture allows data to move efficiently without hitting external network bottlenecks, a critical advantage for latency-sensitive applications.

Source: Tom's Hardware

Source: Tom's Hardware

Market Implications and Deployment Timeline

Cerebras plans to install AMD Helios systems in its own data centers and integrate them with its WSE racks, with the combined offering scheduled to become available initially through Cerebras Cloud in the second half of 2026

1

. Cerebras shares gained about 4% on Thursday following the announcement

3

. The partnership comes shortly after AMD announced Anthropic as a major customer, confirming a deal to supply up to 2 gigawatts of computing power with the first gigawatt expected online in 2027, alongside an investment of up to $5 billion in Anthropic

4

. Lisa Su suggested this may not be AMD's last collaboration in this space, stating during a press conference that the company expects to pursue more workload disaggregation going forward and will work with multiple companies that have technology that could prove useful

2

. Su also projected that AI could help expand the global computing market to $2 trillion by 2030

4

, underscoring the strategic importance of capturing AI compute resources market share. The AMD-Cerebras collaboration offers a rack-scale product that is more versatile than the relatively rigid LPUs within Nvidia's Groq 3 LPX, which features 256 interconnected Groq 3 LPU accelerators

5

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved