AMD and Cerebras forge strategic partnership to challenge Nvidia's $20 billion Groq acquisition

Reviewed byNidhi Govil

10 Sources

Share

AMD and Cerebras Systems announced a strategic partnership to develop a disaggregated AI inference platform that combines AMD's Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine. The collaboration promises up to 5x higher tokens per second per watt by splitting inference workloads between AMD's EPYC processors and Instinct GPUs for prompt processing, while Cerebras handles memory-intensive token generation with its SRAM-powered accelerators.

News article

AMD and Cerebras Unite to Split AI Inference Workloads

AMD and Cerebras Systems unveiled a strategic partnership on Thursday that pairs AMD's Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine (WSE) to create a disaggregated AI inference platform designed for ultra-low latency performance

1

2

. Announced during AMD CEO Lisa Su's Advancing AI keynote in San Francisco, the collaboration addresses a critical gap in AMD's portfolio that cost Nvidia $20 billion to fill through its acquisition of Groq in December

2

3

. The move reflects how workload disaggregation has become essential for AI hardware providers seeking to optimize different stages of inference processing.

How the Disaggregated AI Inference Platform Works

The new platform assigns specific inference tasks to architectures optimized for them, with AMD's Helios racks featuring EPYC processors and Instinct GPUs handling the compute-intensive prompt processing and large context windows, while Cerebras' WSE takes over the memory-bandwidth-intensive token generation stage

1

5

. Unlike Nvidia's approach with Groq LPUs, which uses SRAM for the prefill stage, AMD and Cerebras invert the specialization: AMD handles prefill while Cerebras manages decode

2

. Cerebras CEO Andrew Feldman emphasized the complementary nature of the technologies, stating that AMD's Instinct and Helios rack lead in performance and memory capacity while Cerebras' Wafer-Scale Engine dominates in SRAM and memory bandwidth

2

. The combined system promises up to 5x higher tokens per second per watt compared to standalone configurations, though this metric was tested against a Cerebras WSE baseline using the open-source Kimi 2.6 1T model

1

4

.

Strategic Implications for AMD's AI Hardware Portfolio

This partnership positions AMD to compete more aggressively in the AI inference market without the massive capital outlay Nvidia committed to access similar SRAM decode technology from Groq

4

. Lisa Su indicated during a press conference that this represents just the beginning of AMD's workload disaggregation strategy, noting that the company will work with multiple partners offering specialized technologies that integrate well with Helios

2

5

. The announcement follows AMD's recent deal with Anthropic to supply up to 2 gigawatts of computing power, with the first gigawatt expected online in 2027, alongside a commitment to invest up to $5 billion in the AI company

5

. Su projects that AI could expand the global computing market to $2 trillion by 2030, underscoring the stakes in this competitive landscape

5

.

Cerebras Gains Market Validation as Stock Rises

Cerebras shares gained approximately 4% on Thursday following the partnership announcement, providing a welcome boost to the volatile stock that has fluctuated significantly since its May listing

3

. After going public at $185, Cerebras stock shot up to $386.34 before falling below $161 in late June, with Thursday's gain pushing shares to $219.80

3

. Cerebras plans to install AMD Helios systems in its own data centers and integrate them with WSE racks, with the combined offering scheduled to become available through Cerebras Cloud in the second half of 2026

1

3

. Feldman, who has previously criticized Nvidia as merely an AI arms dealer, highlighted how critical ultra-low latency has become for AI firms, noting that when something is essential, people want to use it quickly

3

. In January, Cerebras secured a deal with OpenAI worth over $10 billion to deliver 750 megawatts of computing power through 2028, demonstrating its growing presence in AI infrastructure

3

.

What This Means for AI Inference Competition

The partnership matters because AI inference—the process that turns trained models into responses—sits at the heart of speed and cost efficiency for everyday AI services

5

. Where Nvidia requires approximately two thousand Groq LPUs worth of SRAM to serve a trillion-parameter model like Kimi K2.5, AMD and Cerebras will need at most a few dozen wafers, though this advantage depends heavily on model architecture

2

. Cerebras' wafer-scale engines rely on on-chip SRAM rather than HBM4, delivering speeds often exceeding 2,000 tokens per second, making Cerebras one of the fastest inference providers globally

2

. However, the absence of financial details and direct performance comparisons to Nvidia's offerings leaves questions about the partnership's economic structure, particularly as concerns about circular financing in AI hardware deals have increased following Nvidia's move to backstop OpenAI's data center purchases

4

. Server buyers will eventually be able to configure AMD systems with Cerebras' wafer-scale chips directly, expanding deployment options beyond Cerebras Cloud

3

. As workload disaggregation becomes standard practice, watch for additional partnerships from AMD targeting specific AI workloads, and for performance benchmarks comparing this approach against Nvidia's integrated Rubin systems when both platforms reach market in late 2026.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved