7 Sources
[1]
AMD and Cerebras partner on low-latency, high-throughput AI inference -- EPYC processors in Helios rack-scale infrastructure paired with Cerebras' Wafer-Scale Engine (WSE) solutions
AMD and Cerebras Systems on Thursday announced plans to develop a platform that would combine AMD's EPYC processors in Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine (WSE) solutions. Together, the new systems promise to combine low latency of AMD's CPUs and Instinct GPUs with high throughput of Cerebras's Wafer Scale Engines (WSE) processors. AMD and Cerebras expect the new inter-rack-scale platform -- based on AMD Helios rack with EPYC CPUs and Instinct MI400-series accelerators inside -- to be responsible for prompt processing and large context windows, whereas Cerebras' WSE will take care of the memory-bandwidth-intensive token-generation stage. AMD and Cerebras expect their disaggregated inference platform to deliver up to 5X higher tokens per second per watt (T/s/W) by assigning different portions of an inference workload to architectures optimized for them. Therefore, AMD Helios provides rack-scale compute capacity and large volumes of complex requests, whereas the Cerebras WSE handles latency-sensitive token generation. The two compute platforms will operate within a single inference workflow, although the companies have not disclosed additional performance data or explained how the systems will be interconnected. The underlying idea of the AMD + Cerebras platform is essentially the same as Nvidia's CPX concept, but AMD and Cerebras assign the specialized hardware to the opposite inference stage. Nvidia's disaggregated design separates inference into context/prefill and generation/decode. The cancelled Rubin CPX GPU with GDDR7 was optimized specifically for the compute-heavy context/prefill stage, while the regular HBM-equipped Rubin GPUs handle the memory-bandwidth-bound generation stage. By contrast, the AMD and Cerebras platform follows the same disaggregation principle, but the specialization is inverted: AMD's Helios platform with Instinct GPUs handles the prefill stage and processes prompts and large context windows, while the Cerebras WSE takes over the decode stage and handles latency-sensitive token generation. Cerebras plans to install AMD Helios systems in its own data centers and integrate them with its WSE racks. The combined offering is scheduled to become available initially through Cerebras Cloud in the second half of 2026, according to the two companies. Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.
[2]
AMD and Cerebras join forces against Nvidia's Groq LPUs
GPUs are great for training, but for inference, you need a heavy dose of speedy memory to churn out the tokens. AMD has tapped Cerebras Systems to develop a disaggregated compute platform combining Instinct GPUs with the chip startup's SRAM-powered AI accelerators. The goal: to deliver ultra-low-latency inference for agentic workloads. The collaboration, announced on stage during AMD CEO Lisa Su's Advancing AI keynote Thursday, closes a gap in AMD's portfolio that cost Nvidia $20 billion to acquihire from Groq back in December. Cerebras CEO and cofounder Andrew Feldman is no fan of Nvidia, having previously denigrated the GPU giant as a mere AI arms dealer. And unlike GPUs, Cerebras' wafer scale engines (WSE) don't rely on HBM4 but instead use on-chip SRAM that's orders of magnitude faster. This has made Cerebras one of the fastest inference providers in the world, with output speeds often exceeding 2,000 tokens a second. By running compute-heavy prompt processing operations on AMD's Instinct GPUs and offloading the memory intensive token generation to Cerebra's WSE accelerator, the duo aims to achieve higher interactivity without compromising on throughput or cost to do it. "What you have with Instinct and the Helios rack is you have the leader in performance and memory capacity. And you marry that with our Wafer Scale Engine, which is the leader in SRAM and in memory bandwidth, and that combination allows us to deliver a solution that is unmatched," Feldman said on stage. Neither company has shared specific figures, but the combination is expected to boost the number of tokens per second generated per watt of electricity consumed by as much as 5x. If any of this sounds familiar, Cerebras' accelerators fill the same role as the Groq 3 LPUs (Language Processing Units) announced alongside Nvidia's Vera Rubin rack systems at GTC in March. But where Nvidia needs two thousand Groq LPUs worth of SRAM to serve a trillion-parameter model like Kimi K2.5, AMD and Cerebras will need at most a few dozen. The combined offering will be available in Cerebras Cloud later this year, but may not be AMD's last deal with the upstart. "There are lots of ways to get workload-specific acceleration done, and I think Cerebras has a very interesting technology. It works very well with Helios," Su said during press conference following the keynote. "The idea of our open ecosystem is frankly that we will work with a number of different companies that may have technology that could be useful." "You can expect that we're going to do more workload disaggregation going forward," she add. ®
[3]
Cerebras stock gains on AMD partnership
Cerebras shares gained about 4% on Thursday after the company forged an agreement with Advanced Micro Devices that involves the two chipmakers to work together on artificial intelligence systems. Cerebras CEO Andrew Feldman said at AMD's AI conference in San Francisco that his company's chips will be used in AMD's Helios AI systems installed in Cerebras data centers starting later this year. Server buyers will be able to configure AMD systems with the company's "wafer-scale" chips as well. At the event, AMD is detailing new chips and its Helios integrated system. The partnership highlights how important "ultra-low latency" has become for AI firms. Chips like those made by Cerebras are configured to provide the first AI answers as quickly as possible, while making tradeoffs in terms of flexibility and total power. The companies claimed that their system would provide five times higher tokens per second per watt than competitors. AMD rival Nvidia bought assets from Groq in December for $20 billion to integrate that company's low-latency technology into its systems. "When something's a necessity, people want to use it, and they want to use it quickly," Feldman said. Cerebras went public in May and has been a particularly volatile stock in its early days. After going public at $185, the stock shot up as high as $386.34 in its debut before falling below $161 in late June. With Thursday's pop, the shares are trading at $219.80. In January, Cerebras announced a deal with OpenAI to deliver 750 megawatts of computing power through 2028, a deal worth over $10 billion.
[4]
AMD and Cerebras join forces on AI inference
Why it matters: Inference is the part of AI computing that turns a trained model into a response, making it central to the speed and cost of everyday AI services. The big picture: The deal comes shortly after AMD announced Anthropic as a major customer, as demand for AI compute continues to grow. Driving the news: Cerebras plans to deploy AMD Helios systems in its data centers, with the joint offering available through Cerebras Cloud later this year. * AMD chips will handle prompt processing and large context windows, while Cerebras systems will accelerate token generation, which requires substantial memory bandwidth. * Earlier this week, AMD confirmed a deal to supply up to 2 gigawatts of computing power to Anthropic, with the first gigawatt expected to come online in 2027. AMD also agreed to invest up to $5 billion in Anthropic. What they're saying: AMD CEO Lisa Su said the partnership reflects a broader shift toward using different chips for different stages of an AI workload. * "I think we're going to see more workload disaggregation," Su said during a briefing with reporters. By the numbers: Su said AI could help expand the global computing market to $2 trillion by 2030.
[5]
AMD Fires Back At NVIDIA's Groq Bet, Fuses The Cerebras Wafer-Scale Engine With Helios For 5x Higher Tokens Per Second Per Watt
NVIDIA scooped up Groq as soon as its LPU showed promise in terms of efficient inferencing. Now, AMD has countered NVIDIA's gambit by partnering with Cerebras to integrate its Helios rack-scale solution with Cerebras' Wafer-Scale Engine, dramatically increasing the inferencing capabilities of the integrated system. LPU vs. Wafer-Scale Engine For the benefit of those who might not be aware, Groq's Language Processing Unit (LPU) clusters hundreds or even thousands of specialized chips together, where each individual chip contains giant blocks of Matrix Multiply (MXM) and Vector (VXM) units as well as around 230MB of blazing-fast SRAM. Also, AI model weights are hard-baked directly into the SRAM, completely bypassing the concept of a memory cache. Crucially, the LPU has no branch predictors or hardware schedulers. Instead, the Groq compiler plans every single calculation down to the exact nanosecond, ensuring that relevant data arrives from the SRAM for processing in a continuous, meticulously planned operational cadence, resulting in extremely fast inferencing. In contrast, Cerebras' Wafer-Scale Engine places an entire AI supercomputer's worth of memory and compute onto a single, giant, interconnected sheet of silicon, where hundreds of thousands of compute cores and tens of GBs of SRAM connect seamlessly, allowing data to move efficiently without ever hitting external network bottlenecks. The LPU uses small isolated pools of the SRAM, where a given AI model is broken apart and then distributed across a large number of specialized chips. Cerebras' Wafer-Scale Engine, however, can hold an entire medium-sized model - or huge pieces of a large model - within its unified chunk of SRAM. Also, its versatility enables training as well as inferencing workloads. AMD is integrating its rack-scale Helios offering with Cerebras' Wafer-Scale Engine to deliver "the ultra-low latency required for the most advanced AI applications" Under AMD's envisioned roadmap, Helios will provide a high-performance, scalable throughput engine, while Cerebras' Wafer-Scale Engine technology will provide ultra-fast, ultra-low latency decode and token generation. Together, the two compute engines are expected to deliver up to 5x higher tokens per second per watt. This arrangement balances the high token generation requirements of volume-based workloads with faster response times prized by coding and other agentic tasks. AMD goes on to note: "Helios provides ultra-high throughput, processing prompts and large context windows. The Cerebras Wafer-Scale Engine accelerates the memory-bandwidth-intensive token generation, with ultra-low latency. By connecting these best-in-class engines through one integrated workflow, the companies are creating a differentiated platform for ultra-low-latency inference without sacrificing throughput or scale." Of course, NVIDIA is now selling its LPU-based rack-scale offering, dubbed the Groq 3 LPX, featuring 256 interconnected Groq 3 LPU accelerators, full liquid cooling, and 315 PFLOPS of inference power. However, the AMD-Cerebras collaboration offers a rack-scale product that is much more versatile than the relatively rigid LPUs within the Groq 3 LPX. Follow Wccftech on Google to get more of our news coverage in your feeds.
[6]
AMD, Cerebras partner on AI inference solution By Investing.com
SAN FRANCISCO and SUNNYVALE, Calif. - AMD (NASDAQ:AMD) and Cerebras Systems (NASDAQ:CBRS) announced today a technical partnership to deliver a disaggregated AI inference solution combining AMD Helios rackscale solutions with the Cerebras Wafer-Scale Engine. The companies unveiled the solution at Advancing AI 2026. The joint offering will deploy AMD Helios alongside Cerebras Wafer-Scale Engine technology in a single inference workflow, according to a press release statement. AMD Helios will provide high-performance, scalable throughput processing, while Cerebras Wafer-Scale Engine technology will handle token generation. The companies stated the combined system is expected to deliver up to 5x higher tokens per second per watt compared to a Cerebras WSE-only configuration, based on modeling conducted in July 2026. The solution addresses different requirements across AI inference workloads. AMD Helios processes prompts and large context windows, while the Cerebras Wafer-Scale Engine handles token generation for applications requiring faster response times. "AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach," said Dr. Lisa Su, chair and CEO of AMD. "Together with Cerebras, we are extending that leadership into the most latency-sensitive applications." Andrew Feldman, CEO and co-founder of Cerebras, said, "Partnering with AMD gives us an incredible opportunity to bring that performance to even more customers." Cerebras plans to deploy AMD Helios systems in its data centers. The joint solution is expected to become available initially through Cerebras Cloud in the second half of 2026. The partnership targets applications including software development, autonomous agents, robotics and scientific discovery where response time affects user experience. This article was generated with the support of AI and reviewed by an editor. For more information see our T&C.
[7]
AMD, Cerebras Team Up on AI Inference Solution
Advanced Micro Devices and Cerebras are partnering on an artificial-intelligence inference offering that aims to deliver the low latency required by advanced AI applications while boosting efficiency. The joint solution combines AMD's Helios rackscale AI infrastructure solutions with Cerebras's Wafer-Scale Engine technology, integrated in a single inference workflow, the companies said Thursday. The solution aims to address demand for infrastructure that matches compute technologies to specific workload requirements, the companies said. AI inference workloads increasingly have different requirements, given high-volume workloads prioritize maximizing token generation, while coding, real-time copilots and live agents demand faster response times. "AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach," AMD Chief Executive Lisa Su said. As part of the collaboration, Cerebras plans to deploy AMD Helios systems in its data centers. The companies expect their joint solution to become available initially through Cerebras Cloud in the second half of this year.
Share
Copy Link
AMD and Cerebras Systems announced a partnership to develop a disaggregated AI inference platform combining AMD's EPYC processors in Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine solutions. The collaboration promises up to 5x higher tokens per second per watt, positioning AMD to compete directly with Nvidia's $20 billion Groq acquisition while addressing the growing demand for ultra-low latency in AI applications.
AMD and Cerebras Systems unveiled plans to develop a disaggregated AI inference platform that combines AMD's EPYC processors in Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine solutions
1
. The partnership, announced during AMD CEO Lisa Su's Advancing AI keynote Thursday, directly addresses a critical gap in AMD's portfolio and positions the company to compete with Nvidia's $20 billion acquisition of Groq in December2
. The collaboration focuses on delivering ultra-low latency AI performance specifically designed for agentic workloads, where speed and responsiveness are paramount.
Source: Axios
The new disaggregated AI inference platform assigns different portions of AI inference workloads to architectures optimized for them, promising up to 5x higher tokens per second per watt
1
. AMD Helios systems with Instinct GPUs handle the compute-heavy prompt processing stage and large context windows, while the Cerebras Wafer-Scale Engine takes over the memory-bandwidth-intensive token generation stage1
. This approach inverts Nvidia's disaggregation strategy, where the cancelled Rubin CPX GPU was optimized for context/prefill while HBM-equipped GPUs handled generation1
.Cerebras CEO and cofounder Andrew Feldman emphasized the technological advantage of combining AMD's leadership in performance and memory capacity with Cerebras' dominance in SRAM and memory bandwidth
2
. Unlike GPUs that rely on HBM4, Cerebras' Wafer-Scale Engine uses on-chip SRAM that's orders of magnitude faster, enabling the company to achieve output speeds often exceeding 2,000 tokens per second2
. The Wafer-Scale Engine places an entire AI supercomputer's worth of memory and compute onto a single, giant, interconnected sheet of silicon, where hundreds of thousands of compute cores and tens of gigabytes of SRAM connect seamlessly5
. This architecture allows data to move efficiently without hitting external network bottlenecks, a critical advantage for latency-sensitive applications.
Source: Tom's Hardware
Related Stories
Cerebras plans to install AMD Helios systems in its own data centers and integrate them with its WSE racks, with the combined offering scheduled to become available initially through Cerebras Cloud in the second half of 2026
1
. Cerebras shares gained about 4% on Thursday following the announcement3
. The partnership comes shortly after AMD announced Anthropic as a major customer, confirming a deal to supply up to 2 gigawatts of computing power with the first gigawatt expected online in 2027, alongside an investment of up to $5 billion in Anthropic4
. Lisa Su suggested this may not be AMD's last collaboration in this space, stating during a press conference that the company expects to pursue more workload disaggregation going forward and will work with multiple companies that have technology that could prove useful2
. Su also projected that AI could help expand the global computing market to $2 trillion by 20304
, underscoring the strategic importance of capturing AI compute resources market share. The AMD-Cerebras collaboration offers a rack-scale product that is more versatile than the relatively rigid LPUs within Nvidia's Groq 3 LPX, which features 256 interconnected Groq 3 LPU accelerators5
.Summarized by
Navi
[2]
1
Technology

2
Policy and Regulation

3
Science and Research
