16 Sources
[1]
Cerebras CS-4 rack systems juice chips for every last drop of AI performance
If high-speed AI inference is what you're after, memory bandwidth is the bottleneck to beat. At a mind-numbing 21.6 petabytes per second (PB/s) of memory bandwidth, Cerebras' dinner-plate-sized AI accelerators were already 1,000x faster than Nvidia's or AMD's best GPUs. The chip newcomer unveiled
[2]
Cerebras launches new server chip and system designed to speed AI chatbots
SAN FRANCISCO, Aug 18 (Reuters) - Cerebras Systems (CBRS.O), opens new tab announced on Tuesday a new version of its server hardware that includes its dinner-plate-sized chips that it says will speed AI chatbot queries. Cerebras makes AI hardware and chips that compete with Nvidia (NVDA.O), opens
[3]
Cerebras launches the CS-4, its first multi-wafer system, though the chip inside is not new
Cerebras has put three of its dinner-plate-sized processors into a single rack for the first time. The CS-4, unveiled on Tuesday at the company's Supernova event and shipping this quarter, is pitched as an inference machine for frontier models, and Cerebras says it runs them up to 30 times faster
[4]
Cerebras CS-4 server system claims 30x faster AI inference
Cerebras Systems introduced the CS-4 on Tuesday, a new rack-scale server system the company says delivers up to 30 times faster AI inference than GPU-based systems and twice the speed of its predecessor, the CS-3. The CS-4 is built around three of Cerebras' newly announced Wafer Scale Engine 3
[5]
Cerebras launches CS-4 AI accelerator, boasting 30x speed over GPUs
Cerebras Systems unveiled the CS-4, a next-generation rack-scale AI accelerator, at its Supernova event in San Francisco on August 19, 2026. The company claims the CS-4 delivers up to 30 times faster inference than GPU-based solutions, positioning it as a significant competitor to Nvidia in the AI
[6]
Cerebras Systems: Cerebras launches new server chip and system designed to speed AI chatbots
Cerebras makes AI hardware and chips that compete with Nvidia and targets the portion of AI called inference, the computing process of generating an answer in a chatbot such as Anthropic's Claude. Cerebras Systems announced on Tuesday a new version of its server hardware that includes its
[7]
Cerebras CS-4 Generates In 1 Second What A GPU Rack Needs 30 Seconds For, Powered By 4-Trillion-Transistor WSE-3 Turbo
Cerebras, the creators of the wafer-scale engine chip, have released their latest CS-4 solution that packs its brand-new WSE-3 Turbo chip. Cerebras CS-4 Is Powered By Wafer-Scale Chips, Packing Up To 250 PFLOPS Per Wafer & Double The Bandwidth of 43.2 PB/s Using 44 GB SRAM Back in 2024, Cerebras
[8]
Cerebras Extends the Lead In Fast Inference Space with New Launch: Analyst - Cerebras Systems (NASDAQ:CBR
Cerebras Systems Inc. (NASDAQ:CBRS) on Tuesday unveiled the CS-4, a rack-scale AI system powered by three of its Wafer Scale Engine-3 Turbo chips and designed to speed up response generation from AI models. CEO Andrew Feldman said Cerebras expects to deliver 600 megawatts of computing capacity by
[9]
Cerebras CS-4: Up To 30X Faster Than Nvidia GPUs - Cerebras Systems (NASDAQ:CBRS)
Cerebras Systems Inc. (NASDAQ:CBRS) unveiled a new AI system Tuesday that it says can generate answers up to 30 times faster than GPU-based alternatives, taking another shot at Nvidia Corp. (NASDAQ:NVDA) in the fast-growing market for AI inference. The new CS-4 hit speeds of more than 4,400 tokens
[10]
Cerebras's New Server Chip and Rack System Aims to Deliver Max Speed to AI Chatbots
They aim to deliver 600 MW worth of computing power by end-2027 with plans focused on speeding up the amount of data that its future chips and systems can crunch Chip newcomer Cerebras Systems has announced a new version of its server hardware that includes dinner-place sized chips aimed to speed
[11]
Cerebras powers OpenAI's GPT-5.6 Sol ultrafast mode at 750 tokens/sec By Investing.com
SUNNYVALE, Calif. - Cerebras Systems (NASDAQ:CBRS) announced today that it is powering Ultrafast mode, a new service tier in the OpenAI API for GPT-5.6 Sol, according to a press release statement. The service, available initially in limited preview to OpenAI customers, runs GPT-5.6 Sol at up to
[12]
Cerebras' CS-4 Launch Reinforces Thesis About Differentiation Company Architecture Offers for Fast Inference, UBS Says
Cerebras Systems Inc. is an artificial intelligence (AI) infrastructure company that designs and manufactures an AI compute platform comprised of proprietary systems and software. The Company's products include inference Cloud, Training Cloud, CS-3 system, AI supercomputer, Wafer Scale Engine and
[13]
Cerebras Systems Unveils CS-4 AI Accelerator
Cerebras Systems introduced the Cerebras CS-4, the fastest AI accelerator in the industry. The CS-4 is a rack-scale solution built from three new Wafer Scale Engines and revolutionary rack and system designs. The CS-4 is the first member of the next-generation Cerebras Nexus rack-scale platform
[14]
Cerebras Systems Launches New CS-4 AI Accelerator
Cerebras Systems Inc. is an artificial intelligence (AI) infrastructure company that designs and manufactures an AI compute platform comprised of proprietary systems and software. The Company's products include inference Cloud, Training Cloud, CS-3 system, AI supercomputer, Wafer Scale Engine and
[15]
Cerebras launches new server chip and system designed to speed AI chatbots
Cerebras Systems Inc. is an artificial intelligence (AI) infrastructure company that designs and manufactures an AI compute platform comprised of proprietary systems and software. The Company's products include inference Cloud, Training Cloud, CS-3 system, AI supercomputer, Wafer Scale Engine and
[16]
Cerebras Systems Powers OpenAI's GPT-5.6 Sol Ultrafast Mode
Cerebras Systems Inc. is an artificial intelligence (AI) infrastructure company that designs and manufactures an AI compute platform comprised of proprietary systems and software. The Company's products include inference Cloud, Training Cloud, CS-3 system, AI supercomputer, Wafer Scale Engine and
Share
Copy Link
Cerebras launched the CS-4, its first multi-wafer AI inference system, at the Supernova event in San Francisco. The rack houses three WSE-3 Turbo processors delivering 750 petaflops of compute and claims 30x faster inference than GPU-based systems. First shipments begin this quarter as the company targets 600 megawatts of computing power by end of 2027.

Cerebras Systems introduced the Cerebras CS-4 at its Supernova event in San Francisco on August 19, 2026, marking the company's first multi-wafer AI inference system
3
. The next-generation rack-scale AI accelerator houses three WSE-3 Turbo processors and ships this quarter, five days after OpenAI's Ultrafast mode running GPT-5.6 Sol roughly 14 times faster on Cerebras silicon went live3
. This launch represents the first hardware Cerebras has shipped since its $5.55 billion Nasdaq debut in May, the largest US tech listing since Snowflake3
.The high-performance rack-scale server system delivers up to 30x faster AI inference than GPU-based systems and twice the speed of its predecessor, the CS-3
4
. Built on the company's Nexus rack systems platform architecture, the CS-4 is designed to speed AI chatbot inference queries and handle frontier models2
3
.Each CS-4 carries three WSE-3 Turbo wafers for a combined 750 petaflops of sparse FP16 compute, 129.6 petabytes per second of memory bandwidth, and 7.2 terabits per second of I/O bandwidth
3
4
. The Wafer Scale Engine 3 Turbo promises twice the compute, memory fabric, and I/O bandwidth of the two-year-old WSE-31
.What makes this achievement notable is that the WSE-3 Turbo isn't new silicon. The chip carries the same four trillion transistors, the same 900,000 cores, the same 44GB of on-chip SRAM, and the same TSMC 5-nanometer manufacturing process as the WSE-3 it replaces
1
2
3
. Cerebras accomplishes the performance gains by pushing its existing Wafer Scale Engine harder through improved power delivery, enabling higher operating frequencies and faster token generation1
.Estimates suggest Cerebras is now running the silicon at 2.8 GHz, up from 1.4 GHz last generation
1
3
. This clock bump delivers the claimed tenfold gain in throughput per watt over the previous generation1
. Each WSE-3 Turbo boasts 250 petaflops of AI compute and 43.2 petabytes per second of memory bandwidth1
5
.The CS-4 introduces a modular "backpack" architecture that reduces the number of components by 50 percent compared to its predecessor
3
5
. Chief Technology Officer Sean Lie said at a media briefing in San Francisco this design would speed data center construction2
.Cerebras' chips are now housed in self-contained systems with all control electronics on board that plug into the back of the rack
1
. The front of the rack is dedicated to power shelves, with power conversion moved a hundred times closer to the processors3
. The system features switchless chip-to-chip connections with a 2D torus topology, lowering interconnect latency from five microseconds to two5
.Estimates put the CS-4 at 120 to 140 kilowatts per rack, roughly half what comparable AMD and Nvidia rack systems draw, which supports the claimed tenfold gain in throughput per watt
3
.Related Stories
In a benchmark on GPT-OSS-120B, a 120-billion-parameter model, the system produced more than 4,400 tokens per second per user
4
5
. This stands in stark contrast to approximately 350 tokens per second on the fastest GPU-based systems5
. The system supports models above 50 trillion parameters with wafer-to-wafer latency falling to two microseconds from five3
.Cerebras has partnered with AWS and AMD to offload compute-intensive prompt processing to their respective Trainium XPUs and Instinct GPUs in a disaggregated inference pipeline
1
. For AI inference, Cerebras' chips now function primarily as decode accelerators, similar to how Nvidia uses Groq LPUs in its LPX rack systems1
.CEO Andrew Feldman emphasized that "in AI, speed is productivity" and told Reuters the company expects to get "four times as fast between now and the end of 2027, and 20 times more throughput"
2
3
. Cerebras expects to deliver 600 megawatts of computing power by the end of 20272
5
.Last week, Cerebras reported an adjusted loss of $6.9 million on sales of $180.1 million
2
5
. Second-quarter revenue came in at $180.1 million, up 74 percent year on year with cloud revenue nearly quadrupling, but it fell sequentially from $193.4 million in the first quarter3
. The quarter produced a GAAP net loss of $450.5 million and $25.4 billion in remaining performance obligations3
.Cerebras guides to $880 million to $890 million of full-year revenue and 600 megawatts of data center capacity live and under contract by the end of 2027
3
. G42 and the Mohamed bin Zayed University of Artificial Intelligence together accounted for around 86 percent of 2025 revenue, though second-quarter additions included Cognition, Lovable, CrowdStrike, Block, and Figma3
. The OpenAI contract signed in January, worth more than $10 billion at signature, represents the company's answer to customer concentration concerns3
.Cerebras plans another generation of the chip and server in 2027
2
. The harder test comes when the next generation will need to arrive on new silicon rather than a faster clock, particularly as the company prioritizes SRAM capacity over compute for agentic AI applications where speed enables "more than an order of magnitude as much reasoning, verification, or tool use," according to CTO Sean Lie3
.Summarized by
Navi
[3]
12 Mar 2025•Technology

08 Jun 2026•Business and Economy

31 Jan 2025•Technology

1
Technology

2
Policy and Regulation

3
Technology
