10 Sources
[1]
AMD and Cerebras partner on low-latency, high-throughput AI inference -- EPYC processors in Helios rack-scale infrastructure paired with Cerebras' Wafer-Scale Engine (WSE) solutions
AMD and Cerebras Systems on Thursday announced plans to develop a platform that would combine AMD's EPYC processors in Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine (WSE) solutions. Together, the new systems promise to combine low latency of AMD's CPUs and Instinct GPUs with high throughput of Cerebras's Wafer Scale Engines (WSE) processors. AMD and Cerebras expect the new inter-rack-scale platform -- based on AMD Helios rack with EPYC CPUs and Instinct MI400-series accelerators inside -- to be responsible for prompt processing and large context windows, whereas Cerebras' WSE will take care of the memory-bandwidth-intensive token-generation stage. AMD and Cerebras expect their disaggregated inference platform to deliver up to 5X higher tokens per second per watt (T/s/W) by assigning different portions of an inference workload to architectures optimized for them. Therefore, AMD Helios provides rack-scale compute capacity and large volumes of complex requests, whereas the Cerebras WSE handles latency-sensitive token generation. The two compute platforms will operate within a single inference workflow, although the companies have not disclosed additional performance data or explained how the systems will be interconnected. The underlying idea of the AMD + Cerebras platform is essentially the same as Nvidia's CPX concept, but AMD and Cerebras assign the specialized hardware to the opposite inference stage. Nvidia's disaggregated design separates inference into context/prefill and generation/decode. The cancelled Rubin CPX GPU with GDDR7 was optimized specifically for the compute-heavy context/prefill stage, while the regular HBM-equipped Rubin GPUs handle the memory-bandwidth-bound generation stage. By contrast, the AMD and Cerebras platform follows the same disaggregation principle, but the specialization is inverted: AMD's Helios platform with Instinct GPUs handles the prefill stage and processes prompts and large context windows, while the Cerebras WSE takes over the decode stage and handles latency-sensitive token generation. Cerebras plans to install AMD Helios systems in its own data centers and integrate them with its WSE racks. The combined offering is scheduled to become available initially through Cerebras Cloud in the second half of 2026, according to the two companies. Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.
[2]
AMD and Cerebras join forces against Nvidia's Groq LPUs
GPUs are great for training, but for inference, you need a heavy dose of speedy memory to churn out the tokens. AMD has tapped Cerebras Systems to develop a disaggregated compute platform combining Instinct GPUs with the chip startup's SRAM-powered AI accelerators. The goal: to deliver ultra-low-latency inference for agentic workloads. The collaboration, announced on stage during AMD CEO Lisa Su's Advancing AI keynote Thursday, closes a gap in AMD's portfolio that cost Nvidia $20 billion to acquihire from Groq back in December. Cerebras CEO and cofounder Andrew Feldman is no fan of Nvidia, having previously denigrated the GPU giant as a mere AI arms dealer. And unlike GPUs, Cerebras' wafer scale engines (WSE) don't rely on HBM4 but instead use on-chip SRAM that's orders of magnitude faster. This has made Cerebras one of the fastest inference providers in the world, with output speeds often exceeding 2,000 tokens a second. By running compute-heavy prompt processing operations on AMD's Instinct GPUs and offloading the memory intensive token generation to Cerebra's WSE accelerator, the duo aims to achieve higher interactivity without compromising on throughput or cost to do it. "What you have with Instinct and the Helios rack is you have the leader in performance and memory capacity. And you marry that with our Wafer Scale Engine, which is the leader in SRAM and in memory bandwidth, and that combination allows us to deliver a solution that is unmatched," Feldman said on stage. Neither company has shared specific figures, but the combination is expected to boost the number of tokens per second generated per watt of electricity consumed by as much as 5x. If any of this sounds familiar, Cerebras' accelerators fill the same role as the Groq 3 LPUs (Language Processing Units) announced alongside Nvidia's Vera Rubin rack systems at GTC in March. But where Nvidia needs two thousand Groq LPUs worth of SRAM to serve a trillion-parameter model like Kimi K2.5, AMD and Cerebras will need at most a few dozen. The combined offering will be available in Cerebras Cloud later this year, but may not be AMD's last deal with the upstart. "There are lots of ways to get workload-specific acceleration done, and I think Cerebras has a very interesting technology. It works very well with Helios," Su said during press conference following the keynote. "The idea of our open ecosystem is frankly that we will work with a number of different companies that may have technology that could be useful." "You can expect that we're going to do more workload disaggregation going forward," she add. ®
[3]
Cerebras stock gains on AMD partnership
Cerebras shares gained about 4% on Thursday after the company forged an agreement with Advanced Micro Devices that involves the two chipmakers to work together on artificial intelligence systems. Cerebras CEO Andrew Feldman said at AMD's AI conference in San Francisco that his company's chips will be used in AMD's Helios AI systems installed in Cerebras data centers starting later this year. Server buyers will be able to configure AMD systems with the company's "wafer-scale" chips as well. At the event, AMD is detailing new chips and its Helios integrated system. The partnership highlights how important "ultra-low latency" has become for AI firms. Chips like those made by Cerebras are configured to provide the first AI answers as quickly as possible, while making tradeoffs in terms of flexibility and total power. The companies claimed that their system would provide five times higher tokens per second per watt than competitors. AMD rival Nvidia bought assets from Groq in December for $20 billion to integrate that company's low-latency technology into its systems. "When something's a necessity, people want to use it, and they want to use it quickly," Feldman said. Cerebras went public in May and has been a particularly volatile stock in its early days. After going public at $185, the stock shot up as high as $386.34 in its debut before falling below $161 in late June. With Thursday's pop, the shares are trading at $219.80. In January, Cerebras announced a deal with OpenAI to deliver 750 megawatts of computing power through 2028, a deal worth over $10 billion.
[4]
Nvidia paid $20 billion for SRAM decode - AMD just partnered for it instead
* AMD and Cerebras will split inference across two machines, with Helios racks handling prompt processing and the Wafer-Scale Engine generating tokens, available through Cerebras Cloud in H2 2026 * Nvidia is also doing something similar by licensing AI chip startup Groq's SRAM decode technology for $20 billion * The move sees AMD and Cerebras claim 5x higher tokens per watt versus a standalone Cerebras WSE configuration AMD and Cerebras Systems have announced a technical partnership which pairs the former's Helios rackscale system with the latter's Wafer-Scale Engine in what both companies call a disaggregated inference solution. The move has enabled a combined AMD Helios and Cerebras WSE configuration to deliver up to five times the tokens per second per watt (TPS/W) in internal testing by both chip designers. The move aims to address a Cerebras WSE efficiency challenge by offloading prompt processing to AMD's rackscale offering. An efficiency gains-centric play? Both AMD and Cerebras Systems are painting the news as a win, and it very well might be, given the latter's efficiency gains in play and the former's ability to get access to SRAM decode technology without spending the $20 billion Nvidia shelled out at the end of last year for a non-exclusive deal. It must, however, be noted that the efficiency claims of 5 tokens per second per watt are compared against an existing Cerebras WSE (Wafer-Scale Engine) as the baseline, while running the open-source Kimi 2.6 1T model, making them impressive, but without a direct comparison to figures for an Nvidia rackscale offering, one that lacks context, especially when efficiency is the metric. The idea itself is sound and well established in the industry, with WSE known to struggle with the 'prefill' part of the equation while handling the 'decode' segment relatively well, essentially substituting AMD's hardware where Cerebras' equipment falls short. The choice of Kimi 2.6, however, deserves a second look. Moonshot AI's model, released on 20 April 2026, is a mixture-of-experts design with one trillion total parameters but only 32 billion active per token, and it ships natively in INT4. At INT4, the full weight set runs to roughly 500 GB. A single Cerebras wafer holds 44 GB. Even before KV cache, a Cerebras-only deployment needs somewhere north of a dozen wafers just to hold the model, while one Helios rack could hold it around sixty times over. That asymmetry means the five-times figure is measured on a model that is close to the least favorable for a WSE-only configuration. A dense model small enough to sit resident on a handful of wafers could flatter Cerebras considerably more. None of this makes the number wrong, but it does make the case for additional testing to demonstrate both its strengths and weaknesses for different models. A partnership without numbers, for now More importantly, the absence of any financial information might very well be a future story, especially at a time when there are increasing concerns about 'circular financing' in an industry where Nvidia's recent move to backstop OpenAI's data center purchases was seen as a net negative by Wall St, which is already concerned about AI spend and the sustainability of such transactions. AMD has also, in the past (and more recently with Anthropic), linked purchases of its own hardware to investments or stakes it would take in AI companies, moves that the market welcomed earlier but might view with a bit more hostility lately. The announcement comes at a time when Cerebras might need it more than AMD: Cerebras listed on Nasdaq in May, priced at $185, opened at $350, and closed its first day at $311.07 before falling back to around $227 by late June 2026. AMD stock, on the other hand, is up 121.48% year-to-date (YTD) as investors continue to bet heavily on its new Instinct AI processors, and the Cerebras partnership allows it to further consolidate its gains, as this might be seen as another vote of confidence in its current direction by one of its prospective customers. Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.
[5]
AMD and Cerebras join forces on AI inference
Why it matters: Inference is the part of AI computing that turns a trained model into a response, making it central to the speed and cost of everyday AI services. The big picture: The deal comes shortly after AMD announced Anthropic as a major customer, as demand for AI compute continues to grow. Driving the news: Cerebras plans to deploy AMD Helios systems in its data centers, with the joint offering available through Cerebras Cloud later this year. * AMD chips will handle prompt processing and large context windows, while Cerebras systems will accelerate token generation, which requires substantial memory bandwidth. * Earlier this week, AMD confirmed a deal to supply up to 2 gigawatts of computing power to Anthropic, with the first gigawatt expected to come online in 2027. AMD also agreed to invest up to $5 billion in Anthropic. What they're saying: AMD CEO Lisa Su said the partnership reflects a broader shift toward using different chips for different stages of an AI workload. * "I think we're going to see more workload disaggregation," Su said during a briefing with reporters. By the numbers: Su said AI could help expand the global computing market to $2 trillion by 2030.
[6]
Disaggregated AI inference with Cerebras and AMD
Cerebras and AMD partner to build the world's fastest disaggregated AI inference solution Disaggregated AI inference is proving to be more than a complementary answer to the prefill and decode bottleneck slowing enterprise AI at scale, and Cerebras and AMD just announced a partnership to build the fastest version of it in the world. The recent collaboration pairs AMD's Helios rack-scale architecture for the compute-intensive pre-fill phase with the Cerebras Wafer-Scale Engine for ultra-low-latency decode, according to Julie Choi (pictured), chief marketing officer at Cerebras. The resulting combination delivers 5x higher tokens per second per watt compared to existing solutions. Later this year, Cerebras will bring AMD Helios systems into its own data centers to power the pre-fill layer of the production deployment. "Lisa and Andrew both announced how AMD and Cerebras are collaborating on the world's most powerful disaggregated inference solution," Choi said. "It's a one plus one equals five X in this case." Choi spoke with theCUBE's John Furrier at the Neo4j GraphTalk event during an exclusive broadcast on theCUBE, SiliconANGLE Media's livestreaming studio. They discussed the technical architecture behind disaggregated AI inference, along with what workloads are driving the fastest demand and why the partnership extends to deploying AMD Helios inside Cerebras data centers before the end of 2026. Disaggregated AI inference pairs AMD Helios with Cerebras for maximum throughput and minimum latency The architectural logic is fairly straightforward. Pre-fill is computationally intensive, but manageable with optimized GPU infrastructure like AMD Helios. Decode is constrained by memory bandwidth. The Cerebras Wafer-Scale Engine carries roughly 2,000 times the memory bandwidth of competing Nvidia GPUs, making it purpose-built for the decode bottleneck that limits large-scale AI inference in production. "On the decode portion, this is a memory bandwidth constrained problem," Choi said. "The Cerebras Wafer-Scale Engine has the largest amount of memory bandwidth. 2,000 times more than Nvidia GPUs." The workloads driving the most demand for disaggregated AI inference are agentic coding, real-time voice and multimodal generation. These are all categories where response speed is a functional requirement. The 5x throughput gain means the same infrastructure can serve dramatically more concurrent users, with joint go-to-market efforts expected before the end of the year, Choi noted. Cerebras will also bring AMD Helios systems into its own data centers this year to power the pre-fill layer, making the partnership both a product collaboration and a production infrastructure commitment. "Our vision is to really provide this speed and max intelligence, no trade-off, to every developer on Earth," Choi said. Here's the complete video interview, part of SiliconANGLE's and theCUBE's coverage of the Neo4j GraphTalk event: (* Disclosure: TheCUBE is a paid media partner for the Neo4j GraphTalk event. Neither Neo4j, the sponsor of theCUBE's event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)
[7]
OpenAI Finds That Systems From AMD's Helios Partner, Cerebras, Are "Incredible" At Inference Tasks, Showing That AMD Chose Wisely
AMD appears to have chosen its partner quite wisely to bolster the Helios rack-scale system's inferencing capabilities, with an OpenAI researcher recently singing what is nothing short a paean in favor of Cerebras' Wafer-Scale Engine. OpenAI researcher on the capabilities of chips from AMD Helios partner, Cerebras: "They're incredible because they have such fast inference" As we explained in a dedicated post recently, Helios is AMD's first full-stack rack-level solution for AI workloads, consisting of: To further bolster the utility of its Helios systems for AI workloads, AMD has partnered with Cerebras, and plans to integrate Cerebras' Wafer-Scale Engine with its Helios systems for blazing-fast inference. For the benefit of those who might not be aware, Cerebras' Wafer-Scale Engine places an entire AI supercomputer's worth of memory and compute onto a single, giant, interconnected sheet of silicon, where hundreds of thousands of compute cores and tens of GBs of SRAM connect seamlessly, allowing data to move efficiently without ever hitting external network bottlenecks. Basically, Cerebras' Wafer-Scale Engine can hold an entire medium-sized model - or huge pieces of a large model - within its unified chunk of SRAM. Also, the versatility of Cerebras' approach enables training as well as inferencing workloads. Under AMD's envisioned roadmap, Helios will provide a high-performance, scalable throughput engine, while Cerebras' technology will provide ultra-fast, ultra-low latency decode and token generation. Together, the two compute engines are expected to deliver up to 5x higher tokens per second per watt. This brings us to the core topic. An OpenAI researcher, Jeffrey Wang, has just generously praised chips from Cerebras that presumably leverage its Wafer-Scale Engine architecture, noting: "Internally, we have some OpenAI models that are on Cerebras chips. They're incredible because they have such fast inference." Wang goes on to declare: ""And what this means for me on the day-to-day is: whereas formerly I might have to wait a couple minutes for a task to finish, it now finishes for me before I even have the opportunity to context-switch. It makes me way more productive." As a refresher, AI models typically context-switch when they jump from one task to another. This could occur within a multi-agentic setting where the model has to perform a series of diverse tasks, or in Mixture of Experts (MoE) models, where various sub-neural networks specialize in performing a particular task. The fact that OpenAI is already impressed with Cerebras' systems suggests that AMD's Helios rack-scale systems - integrated with Wafer-Scale Engines - will likely sell like hotcakes upon debut. This becomes all the more important when you consider that inference costs are now the biggest determinant of data center profitability. Follow Wccftech on Google to get more of our news coverage in your feeds.
[8]
AMD Fires Back At NVIDIA's Groq Bet, Fuses The Cerebras Wafer-Scale Engine With Helios For 5x Higher Tokens Per Second Per Watt
NVIDIA scooped up Groq as soon as its LPU showed promise in terms of efficient inferencing. Now, AMD has countered NVIDIA's gambit by partnering with Cerebras to integrate its Helios rack-scale solution with Cerebras' Wafer-Scale Engine, dramatically increasing the inferencing capabilities of the integrated system. LPU vs. Wafer-Scale Engine For the benefit of those who might not be aware, Groq's Language Processing Unit (LPU) clusters hundreds or even thousands of specialized chips together, where each individual chip contains giant blocks of Matrix Multiply (MXM) and Vector (VXM) units as well as around 230MB of blazing-fast SRAM. Also, AI model weights are hard-baked directly into the SRAM, completely bypassing the concept of a memory cache. Crucially, the LPU has no branch predictors or hardware schedulers. Instead, the Groq compiler plans every single calculation down to the exact nanosecond, ensuring that relevant data arrives from the SRAM for processing in a continuous, meticulously planned operational cadence, resulting in extremely fast inferencing. In contrast, Cerebras' Wafer-Scale Engine places an entire AI supercomputer's worth of memory and compute onto a single, giant, interconnected sheet of silicon, where hundreds of thousands of compute cores and tens of GBs of SRAM connect seamlessly, allowing data to move efficiently without ever hitting external network bottlenecks. The LPU uses small isolated pools of the SRAM, where a given AI model is broken apart and then distributed across a large number of specialized chips. Cerebras' Wafer-Scale Engine, however, can hold an entire medium-sized model - or huge pieces of a large model - within its unified chunk of SRAM. Also, its versatility enables training as well as inferencing workloads. AMD is integrating its rack-scale Helios offering with Cerebras' Wafer-Scale Engine to deliver "the ultra-low latency required for the most advanced AI applications" Under AMD's envisioned roadmap, Helios will provide a high-performance, scalable throughput engine, while Cerebras' Wafer-Scale Engine technology will provide ultra-fast, ultra-low latency decode and token generation. Together, the two compute engines are expected to deliver up to 5x higher tokens per second per watt. This arrangement balances the high token generation requirements of volume-based workloads with faster response times prized by coding and other agentic tasks. AMD goes on to note: "Helios provides ultra-high throughput, processing prompts and large context windows. The Cerebras Wafer-Scale Engine accelerates the memory-bandwidth-intensive token generation, with ultra-low latency. By connecting these best-in-class engines through one integrated workflow, the companies are creating a differentiated platform for ultra-low-latency inference without sacrificing throughput or scale." Of course, NVIDIA is now selling its LPU-based rack-scale offering, dubbed the Groq 3 LPX, featuring 256 interconnected Groq 3 LPU accelerators, full liquid cooling, and 315 PFLOPS of inference power. However, the AMD-Cerebras collaboration offers a rack-scale product that is much more versatile than the relatively rigid LPUs within the Groq 3 LPX. Follow Wccftech on Google to get more of our news coverage in your feeds.
[9]
AMD, Cerebras partner on AI inference solution By Investing.com
SAN FRANCISCO and SUNNYVALE, Calif. - AMD (NASDAQ:AMD) and Cerebras Systems (NASDAQ:CBRS) announced today a technical partnership to deliver a disaggregated AI inference solution combining AMD Helios rackscale solutions with the Cerebras Wafer-Scale Engine. The companies unveiled the solution at Advancing AI 2026. The joint offering will deploy AMD Helios alongside Cerebras Wafer-Scale Engine technology in a single inference workflow, according to a press release statement. AMD Helios will provide high-performance, scalable throughput processing, while Cerebras Wafer-Scale Engine technology will handle token generation. The companies stated the combined system is expected to deliver up to 5x higher tokens per second per watt compared to a Cerebras WSE-only configuration, based on modeling conducted in July 2026. The solution addresses different requirements across AI inference workloads. AMD Helios processes prompts and large context windows, while the Cerebras Wafer-Scale Engine handles token generation for applications requiring faster response times. "AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach," said Dr. Lisa Su, chair and CEO of AMD. "Together with Cerebras, we are extending that leadership into the most latency-sensitive applications." Andrew Feldman, CEO and co-founder of Cerebras, said, "Partnering with AMD gives us an incredible opportunity to bring that performance to even more customers." Cerebras plans to deploy AMD Helios systems in its data centers. The joint solution is expected to become available initially through Cerebras Cloud in the second half of 2026. The partnership targets applications including software development, autonomous agents, robotics and scientific discovery where response time affects user experience. This article was generated with the support of AI and reviewed by an editor. For more information see our T&C.
[10]
AMD, Cerebras Team Up on AI Inference Solution
Advanced Micro Devices and Cerebras are partnering on an artificial-intelligence inference offering that aims to deliver the low latency required by advanced AI applications while boosting efficiency. The joint solution combines AMD's Helios rackscale AI infrastructure solutions with Cerebras's Wafer-Scale Engine technology, integrated in a single inference workflow, the companies said Thursday. The solution aims to address demand for infrastructure that matches compute technologies to specific workload requirements, the companies said. AI inference workloads increasingly have different requirements, given high-volume workloads prioritize maximizing token generation, while coding, real-time copilots and live agents demand faster response times. "AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach," AMD Chief Executive Lisa Su said. As part of the collaboration, Cerebras plans to deploy AMD Helios systems in its data centers. The companies expect their joint solution to become available initially through Cerebras Cloud in the second half of this year.
Share
Copy Link
AMD and Cerebras Systems announced a strategic partnership to develop a disaggregated AI inference platform that combines AMD's Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine. The collaboration promises up to 5x higher tokens per second per watt by splitting inference workloads between AMD's EPYC processors and Instinct GPUs for prompt processing, while Cerebras handles memory-intensive token generation with its SRAM-powered accelerators.

AMD and Cerebras Systems unveiled a strategic partnership on Thursday that pairs AMD's Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine (WSE) to create a disaggregated AI inference platform designed for ultra-low latency performance
1
2
. Announced during AMD CEO Lisa Su's Advancing AI keynote in San Francisco, the collaboration addresses a critical gap in AMD's portfolio that cost Nvidia $20 billion to fill through its acquisition of Groq in December2
3
. The move reflects how workload disaggregation has become essential for AI hardware providers seeking to optimize different stages of inference processing.The new platform assigns specific inference tasks to architectures optimized for them, with AMD's Helios racks featuring EPYC processors and Instinct GPUs handling the compute-intensive prompt processing and large context windows, while Cerebras' WSE takes over the memory-bandwidth-intensive token generation stage
1
5
. Unlike Nvidia's approach with Groq LPUs, which uses SRAM for the prefill stage, AMD and Cerebras invert the specialization: AMD handles prefill while Cerebras manages decode2
. Cerebras CEO Andrew Feldman emphasized the complementary nature of the technologies, stating that AMD's Instinct and Helios rack lead in performance and memory capacity while Cerebras' Wafer-Scale Engine dominates in SRAM and memory bandwidth2
. The combined system promises up to 5x higher tokens per second per watt compared to standalone configurations, though this metric was tested against a Cerebras WSE baseline using the open-source Kimi 2.6 1T model1
4
.This partnership positions AMD to compete more aggressively in the AI inference market without the massive capital outlay Nvidia committed to access similar SRAM decode technology from Groq
4
. Lisa Su indicated during a press conference that this represents just the beginning of AMD's workload disaggregation strategy, noting that the company will work with multiple partners offering specialized technologies that integrate well with Helios2
5
. The announcement follows AMD's recent deal with Anthropic to supply up to 2 gigawatts of computing power, with the first gigawatt expected online in 2027, alongside a commitment to invest up to $5 billion in the AI company5
. Su projects that AI could expand the global computing market to $2 trillion by 2030, underscoring the stakes in this competitive landscape5
.Related Stories
Cerebras shares gained approximately 4% on Thursday following the partnership announcement, providing a welcome boost to the volatile stock that has fluctuated significantly since its May listing
3
. After going public at $185, Cerebras stock shot up to $386.34 before falling below $161 in late June, with Thursday's gain pushing shares to $219.803
. Cerebras plans to install AMD Helios systems in its own data centers and integrate them with WSE racks, with the combined offering scheduled to become available through Cerebras Cloud in the second half of 20261
3
. Feldman, who has previously criticized Nvidia as merely an AI arms dealer, highlighted how critical ultra-low latency has become for AI firms, noting that when something is essential, people want to use it quickly3
. In January, Cerebras secured a deal with OpenAI worth over $10 billion to deliver 750 megawatts of computing power through 2028, demonstrating its growing presence in AI infrastructure3
.The partnership matters because AI inference—the process that turns trained models into responses—sits at the heart of speed and cost efficiency for everyday AI services
5
. Where Nvidia requires approximately two thousand Groq LPUs worth of SRAM to serve a trillion-parameter model like Kimi K2.5, AMD and Cerebras will need at most a few dozen wafers, though this advantage depends heavily on model architecture2
. Cerebras' wafer-scale engines rely on on-chip SRAM rather than HBM4, delivering speeds often exceeding 2,000 tokens per second, making Cerebras one of the fastest inference providers globally2
. However, the absence of financial details and direct performance comparisons to Nvidia's offerings leaves questions about the partnership's economic structure, particularly as concerns about circular financing in AI hardware deals have increased following Nvidia's move to backstop OpenAI's data center purchases4
. Server buyers will eventually be able to configure AMD systems with Cerebras' wafer-scale chips directly, expanding deployment options beyond Cerebras Cloud3
. As workload disaggregation becomes standard practice, watch for additional partnerships from AMD targeting specific AI workloads, and for performance benchmarks comparing this approach against Nvidia's integrated Rubin systems when both platforms reach market in late 2026.Summarized by
Navi
[2]
20 Jul 2026•Technology

12 Mar 2025•Technology

14 Jan 2026•Technology

1
Technology

2
Science and Research

3
Technology

1
AI Agents Escape Safety Tests, Start Turf Wars and Hack Real Systems in Alarming Security Incidents

2
DeepMind's AI weather model gives forecasters an extra day to prepare for deadly tropical cyclones

3
Google Unveils Pixel 11 Series With Gemini AI, New Pixel Tag Tracker and Watch 5 at Made by Google 2026
