10 Sources
[1]
AMD and Cerebras partner on low-latency, high-throughput AI inference -- EPYC processors in Helios rack-scale infrastructure paired with Cerebras' Wafer-Scale Engine (WSE) solutions
AMD and Cerebras Systems on Thursday announced plans to develop a platform that would combine AMD's EPYC processors in Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine (WSE) solutions. Together, the new systems promise to combine low latency of AMD's CPUs and Instinct GPUs with
[2]
AMD and Cerebras join forces against Nvidia's Groq LPUs
GPUs are great for training, but for inference, you need a heavy dose of speedy memory to churn out the tokens. AMD has tapped Cerebras Systems to develop a disaggregated compute platform combining Instinct GPUs with the chip startup's SRAM-powered AI accelerators. The goal: to deliver
[3]
Cerebras stock gains on AMD partnership
Cerebras shares gained about 4% on Thursday after the company forged an agreement with Advanced Micro Devices that involves the two chipmakers to work together on artificial intelligence systems. Cerebras CEO Andrew Feldman said at AMD's AI conference in San Francisco that his company's chips will
[4]
Nvidia paid $20 billion for SRAM decode - AMD just partnered for it instead
* AMD and Cerebras will split inference across two machines, with Helios racks handling prompt processing and the Wafer-Scale Engine generating tokens, available through Cerebras Cloud in H2 2026 * Nvidia is also doing something similar by licensing AI chip startup Groq's SRAM decode technology
[5]
AMD and Cerebras join forces on AI inference
Why it matters: Inference is the part of AI computing that turns a trained model into a response, making it central to the speed and cost of everyday AI services. The big picture: The deal comes shortly after AMD announced Anthropic as a major customer, as demand for AI compute continues to
[6]
Disaggregated AI inference with Cerebras and AMD
Cerebras and AMD partner to build the world's fastest disaggregated AI inference solution Disaggregated AI inference is proving to be more than a complementary answer to the prefill and decode bottleneck slowing enterprise AI at scale, and Cerebras and AMD just announced a partnership to build the
[7]
OpenAI Finds That Systems From AMD's Helios Partner, Cerebras, Are "Incredible" At Inference Tasks, Showing That AMD Chose Wisely
AMD appears to have chosen its partner quite wisely to bolster the Helios rack-scale system's inferencing capabilities, with an OpenAI researcher recently singing what is nothing short a paean in favor of Cerebras' Wafer-Scale Engine. OpenAI researcher on the capabilities of chips from AMD Helios
[8]
AMD Fires Back At NVIDIA's Groq Bet, Fuses The Cerebras Wafer-Scale Engine With Helios For 5x Higher Tokens Per Second Per Watt
NVIDIA scooped up Groq as soon as its LPU showed promise in terms of efficient inferencing. Now, AMD has countered NVIDIA's gambit by partnering with Cerebras to integrate its Helios rack-scale solution with Cerebras' Wafer-Scale Engine, dramatically increasing the inferencing capabilities of the
[9]
AMD, Cerebras partner on AI inference solution By Investing.com
SAN FRANCISCO and SUNNYVALE, Calif. - AMD (NASDAQ:AMD) and Cerebras Systems (NASDAQ:CBRS) announced today a technical partnership to deliver a disaggregated AI inference solution combining AMD Helios rackscale solutions with the Cerebras Wafer-Scale Engine. The companies unveiled the solution at
[10]
AMD, Cerebras Team Up on AI Inference Solution
Advanced Micro Devices and Cerebras are partnering on an artificial-intelligence inference offering that aims to deliver the low latency required by advanced AI applications while boosting efficiency. The joint solution combines AMD's Helios rackscale AI infrastructure solutions with Cerebras's
Share
Copy Link
AMD and Cerebras Systems announced a strategic partnership to develop a disaggregated AI inference platform that combines AMD's Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine. The collaboration promises up to 5x higher tokens per second per watt by splitting inference workloads between AMD's EPYC processors and Instinct GPUs for prompt processing, while Cerebras handles memory-intensive token generation with its SRAM-powered accelerators.

AMD and Cerebras Systems unveiled a strategic partnership on Thursday that pairs AMD's Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine (WSE) to create a disaggregated AI inference platform designed for ultra-low latency performance
1
2
. Announced during AMD CEO Lisa Su's Advancing AI keynote in San Francisco, the collaboration addresses a critical gap in AMD's portfolio that cost Nvidia $20 billion to fill through its acquisition of Groq in December2
3
. The move reflects how workload disaggregation has become essential for AI hardware providers seeking to optimize different stages of inference processing.The new platform assigns specific inference tasks to architectures optimized for them, with AMD's Helios racks featuring EPYC processors and Instinct GPUs handling the compute-intensive prompt processing and large context windows, while Cerebras' WSE takes over the memory-bandwidth-intensive token generation stage
1
5
. Unlike Nvidia's approach with Groq LPUs, which uses SRAM for the prefill stage, AMD and Cerebras invert the specialization: AMD handles prefill while Cerebras manages decode2
. Cerebras CEO Andrew Feldman emphasized the complementary nature of the technologies, stating that AMD's Instinct and Helios rack lead in performance and memory capacity while Cerebras' Wafer-Scale Engine dominates in SRAM and memory bandwidth2
. The combined system promises up to 5x higher tokens per second per watt compared to standalone configurations, though this metric was tested against a Cerebras WSE baseline using the open-source Kimi 2.6 1T model1
4
.This partnership positions AMD to compete more aggressively in the AI inference market without the massive capital outlay Nvidia committed to access similar SRAM decode technology from Groq
4
. Lisa Su indicated during a press conference that this represents just the beginning of AMD's workload disaggregation strategy, noting that the company will work with multiple partners offering specialized technologies that integrate well with Helios2
5
. The announcement follows AMD's recent deal with Anthropic to supply up to 2 gigawatts of computing power, with the first gigawatt expected online in 2027, alongside a commitment to invest up to $5 billion in the AI company5
. Su projects that AI could expand the global computing market to $2 trillion by 2030, underscoring the stakes in this competitive landscape5
.Related Stories
Cerebras shares gained approximately 4% on Thursday following the partnership announcement, providing a welcome boost to the volatile stock that has fluctuated significantly since its May listing
3
. After going public at $185, Cerebras stock shot up to $386.34 before falling below $161 in late June, with Thursday's gain pushing shares to $219.803
. Cerebras plans to install AMD Helios systems in its own data centers and integrate them with WSE racks, with the combined offering scheduled to become available through Cerebras Cloud in the second half of 20261
3
. Feldman, who has previously criticized Nvidia as merely an AI arms dealer, highlighted how critical ultra-low latency has become for AI firms, noting that when something is essential, people want to use it quickly3
. In January, Cerebras secured a deal with OpenAI worth over $10 billion to deliver 750 megawatts of computing power through 2028, demonstrating its growing presence in AI infrastructure3
.The partnership matters because AI inference—the process that turns trained models into responses—sits at the heart of speed and cost efficiency for everyday AI services
5
. Where Nvidia requires approximately two thousand Groq LPUs worth of SRAM to serve a trillion-parameter model like Kimi K2.5, AMD and Cerebras will need at most a few dozen wafers, though this advantage depends heavily on model architecture2
. Cerebras' wafer-scale engines rely on on-chip SRAM rather than HBM4, delivering speeds often exceeding 2,000 tokens per second, making Cerebras one of the fastest inference providers globally2
. However, the absence of financial details and direct performance comparisons to Nvidia's offerings leaves questions about the partnership's economic structure, particularly as concerns about circular financing in AI hardware deals have increased following Nvidia's move to backstop OpenAI's data center purchases4
. Server buyers will eventually be able to configure AMD systems with Cerebras' wafer-scale chips directly, expanding deployment options beyond Cerebras Cloud3
. As workload disaggregation becomes standard practice, watch for additional partnerships from AMD targeting specific AI workloads, and for performance benchmarks comparing this approach against Nvidia's integrated Rubin systems when both platforms reach market in late 2026.Summarized by
Navi
[2]
20 Jul 2026•Technology

12 Mar 2025•Technology

14 Jan 2026•Technology

1
Technology

2
Technology

3
Policy and Regulation
