Cerebras CS-4 Debuts With 30x Faster AI Inference Using Three Wafer-Scale Chips

16 Sources

Share

Cerebras launched the CS-4, its first multi-wafer AI inference system, at the Supernova event in San Francisco. The rack houses three WSE-3 Turbo processors delivering 750 petaflops of compute and claims 30x faster inference than GPU-based systems. First shipments begin this quarter as the company targets 600 megawatts of computing power by end of 2027.

News article

Cerebras Unveils CS-4 Multi-Wafer AI Inference System

Cerebras Systems introduced the Cerebras CS-4 at its Supernova event in San Francisco on August 19, 2026, marking the company's first multi-wafer AI inference system

3

. The next-generation rack-scale AI accelerator houses three WSE-3 Turbo processors and ships this quarter, five days after OpenAI's Ultrafast mode running GPT-5.6 Sol roughly 14 times faster on Cerebras silicon went live

3

. This launch represents the first hardware Cerebras has shipped since its $5.55 billion Nasdaq debut in May, the largest US tech listing since Snowflake

3

.

The high-performance rack-scale server system delivers up to 30x faster AI inference than GPU-based systems and twice the speed of its predecessor, the CS-3

4

. Built on the company's Nexus rack systems platform architecture, the CS-4 is designed to speed AI chatbot inference queries and handle frontier models

2

3

.

WSE-3 Turbo: Doubling Performance Through Clock Speed

Each CS-4 carries three WSE-3 Turbo wafers for a combined 750 petaflops of sparse FP16 compute, 129.6 petabytes per second of memory bandwidth, and 7.2 terabits per second of I/O bandwidth

3

4

. The Wafer Scale Engine 3 Turbo promises twice the compute, memory fabric, and I/O bandwidth of the two-year-old WSE-3

1

.

What makes this achievement notable is that the WSE-3 Turbo isn't new silicon. The chip carries the same four trillion transistors, the same 900,000 cores, the same 44GB of on-chip SRAM, and the same TSMC 5-nanometer manufacturing process as the WSE-3 it replaces

1

2

3

. Cerebras accomplishes the performance gains by pushing its existing Wafer Scale Engine harder through improved power delivery, enabling higher operating frequencies and faster token generation

1

.

Estimates suggest Cerebras is now running the silicon at 2.8 GHz, up from 1.4 GHz last generation

1

3

. This clock bump delivers the claimed tenfold gain in throughput per watt over the previous generation

1

. Each WSE-3 Turbo boasts 250 petaflops of AI compute and 43.2 petabytes per second of memory bandwidth

1

5

.

Modular Architecture Simplifies Deployment

The CS-4 introduces a modular "backpack" architecture that reduces the number of components by 50 percent compared to its predecessor

3

5

. Chief Technology Officer Sean Lie said at a media briefing in San Francisco this design would speed data center construction

2

.

Cerebras' chips are now housed in self-contained systems with all control electronics on board that plug into the back of the rack

1

. The front of the rack is dedicated to power shelves, with power conversion moved a hundred times closer to the processors

3

. The system features switchless chip-to-chip connections with a 2D torus topology, lowering interconnect latency from five microseconds to two

5

.

Estimates put the CS-4 at 120 to 140 kilowatts per rack, roughly half what comparable AMD and Nvidia rack systems draw, which supports the claimed tenfold gain in throughput per watt

3

.

Performance Benchmarks and Strategic Positioning

In a benchmark on GPT-OSS-120B, a 120-billion-parameter model, the system produced more than 4,400 tokens per second per user

4

5

. This stands in stark contrast to approximately 350 tokens per second on the fastest GPU-based systems

5

. The system supports models above 50 trillion parameters with wafer-to-wafer latency falling to two microseconds from five

3

.

Cerebras has partnered with AWS and AMD to offload compute-intensive prompt processing to their respective Trainium XPUs and Instinct GPUs in a disaggregated inference pipeline

1

. For AI inference, Cerebras' chips now function primarily as decode accelerators, similar to how Nvidia uses Groq LPUs in its LPX rack systems

1

.

CEO Andrew Feldman emphasized that "in AI, speed is productivity" and told Reuters the company expects to get "four times as fast between now and the end of 2027, and 20 times more throughput"

2

3

. Cerebras expects to deliver 600 megawatts of computing power by the end of 2027

2

5

.

Financial Performance and Market Position

Last week, Cerebras reported an adjusted loss of $6.9 million on sales of $180.1 million

2

5

. Second-quarter revenue came in at $180.1 million, up 74 percent year on year with cloud revenue nearly quadrupling, but it fell sequentially from $193.4 million in the first quarter

3

. The quarter produced a GAAP net loss of $450.5 million and $25.4 billion in remaining performance obligations

3

.

Cerebras guides to $880 million to $890 million of full-year revenue and 600 megawatts of data center capacity live and under contract by the end of 2027

3

. G42 and the Mohamed bin Zayed University of Artificial Intelligence together accounted for around 86 percent of 2025 revenue, though second-quarter additions included Cognition, Lovable, CrowdStrike, Block, and Figma

3

. The OpenAI contract signed in January, worth more than $10 billion at signature, represents the company's answer to customer concentration concerns

3

.

Cerebras plans another generation of the chip and server in 2027

2

. The harder test comes when the next generation will need to arrive on new silicon rather than a faster clock, particularly as the company prioritizes SRAM capacity over compute for agentic AI applications where speed enables "more than an order of magnitude as much reasoning, verification, or tool use," according to CTO Sean Lie

3

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved