Cerebras Systems partners with Gimlet Labs to deliver ultrafast AI inference at production scale

5 Sources

Share

Cerebras Systems announced a partnership with cloud computing startup Gimlet Labs to deploy 100 megawatts of AI inference capacity. The collaboration will integrate Cerebras CS-4 systems with Gimlet Cloud's disaggregated inference architecture to deliver speeds up to 3,000 tokens per second for demanding agentic and real-time applications.

Cerebras Systems and Gimlet Labs Form Strategic Partnership

Cerebras Systems announced a partnership with cloud computing startup Gimlet Labs to deploy Cerebras-powered AI inference capacity representing roughly 100 megawatts of electrical power

1

2

. The collaboration brings together Cerebras' wafer-scale processors with Gimlet Cloud to create what the companies describe as a purpose-built disaggregated inference architecture spanning datacenter infrastructure to developer APIs

4

. This deal reflects rising interest from AI companies in procuring AI hardware systems capable of speedy AI calculations known as AI inference, the process by which chatbots generate responses to user questions

1

.

Delivering Ultrafast AI Inference Through CS-4 Systems

Cerebras plans to supply its CS-4 systems, unveiled over the summer, to Gimlet Labs over one to two years

1

3

. Gimlet Labs plans to make the CS-4 systems available through its AI cloud in 2027, with the first Cerebras-powered Gimlet Cloud datacenter expected to come online later this year

2

4

. Together, the companies plan to deliver speeds of up to 3,000 tokens per second for demanding agentic applications and real-time applications

2

4

. The companies declined to disclose the financial terms of the deal, though Gimlet Labs will be responsible for maintaining and operating the systems once delivered by Cerebras Systems

1

5

.

Disaggregated Inference Architecture Optimizes Performance

Source: Market Screener

Source: Market Screener

Gimlet Cloud's platform combines Cerebras wafer-scale processors with graphics processing units in what the company calls a disaggregated inference architecture

2

. This approach routes different phases of AI inference—such as prefill, decode, and attention operations—to whichever type of chip is best suited for each task

2

. "By combining Gimlet's multi-silicon software with the Cerebras Wafer Scale Engine, we can run each phase of inference on the hardware best suited to it and plan to deliver up to 3,000 tokens per second at production scale," said Gimlet Labs Co-founder and CEO Zain Asgar

2

. Cerebras Co-founder and CTO Sean Lie noted that combining the company's chips with high-throughput GPUs improves datacenter economics while giving customers a direct path to Cerebras' latest technology

2

.

Target Markets and Use Cases for Fast Inference

Performing inference operations speedily can help in areas such as cybersecurity, voice applications, and financial analysis, according to Gimlet CEO Zain Asgar

1

3

. Gimlet Labs plans to target companies and startups building their entire products around AI, focusing on those requiring the infrastructure to run frontier AI models

1

2

. These frontier models—the most capable AI systems available—demand significant compute resources to run speedily

2

5

. The collaboration builds on joint customer work the two companies have conducted since last year, with an integrated solution already handling inference traffic in private deployments

2

4

.

Cerebras Validates Deployment Flexibility

Source: ET

Source: ET

Cerebras CEO Andrew Feldman described the deal as "further validation from Cerebras Systems of how easy it is to deploy, even in heterogeneous environments," referring to installing the hardware alongside chips and systems from other companies

1

3

. This deployment flexibility matters as companies seek to optimize their AI infrastructure with multiple silicon types. OpenAI and companies such as Nvidia have recognized the importance of speedy inference capabilities

1

3

. Earlier this year, Cerebras Systems signed a deal to supply OpenAI with chips, while Nvidia signed a licensing deal with Groq for chips and hardware last year

1

3

. Cerebras, which reported second-quarter core revenue of $210 million in August and raised its full-year outlook, continues expanding its customer base through similar supply agreements

2

. Gimlet Labs is backed by Andreessen Horowitz and Menlo Ventures

2

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved