Cerebras Systems Powers OpenAI's GPT-5.6 Sol Ultrafast Mode at 750 Tokens Per Second

2 Sources

Share

Cerebras Systems announced it is powering Ultrafast mode for OpenAI's GPT-5.6 Sol, delivering 750 output tokens per second—up to 14 times faster than Standard processing. The service maintains full intelligence while eliminating memory-bandwidth bottlenecks through Cerebras' Wafer-Scale Engine architecture.

Cerebras Systems Launches Ultrafast Mode for OpenAI's GPT-5.6 Sol

Cerebras Systems, an AI infrastructure company specializing in high-performance AI computations, announced it is powering a new Ultrafast mode for OpenAI's GPT-5.6 Sol through the OpenAI API

1

. The service, currently available in limited preview to select OpenAI customers, delivers up to 750 output tokens per second—representing speeds up to 14 times faster than Standard processing while maintaining the same AI model capability and intelligence as GPT-5.6 Sol Standard

1

. Andrew Feldman, CEO and co-founder of Cerebras Systems, stated that "GPT-5.6 Sol on Ultrafast is proof that speed and intelligence are no longer mutually exclusive"

1

. This development marks a shift in how organizations can deploy generative AI applications where reduced latency directly impacts user experience and operational efficiency.

Performance Benchmarks Demonstrate Substantial Speed Advantages

GPT-5.6 Sol Ultrafast significantly outpaces competing models in processing speed. Based on output speeds for Anthropic models reported by Artificial Analysis, the Ultrafast mode runs 5 times faster than Claude Opus 4.8 in Fast mode and 11 times faster than Claude Fable 5

1

. On Humanity's Last Exam, a rigorous 2,500-question benchmark spanning graduate-level subjects, GPT-5.6 Sol Ultrafast completed the entire question set in just over 11 hours, compared to more than three days required by Claude Fable 5

1

. The GDP-Val benchmark, which evaluates knowledge-work tasks, showed Ultrafast delivering a 5.6 times end-to-end speedup

1

. These metrics suggest the technology could transform workflows in research, customer service, and real-time decision-making applications where inference speed directly correlates with business value.

Wafer-Scale Engine Architecture Eliminates Memory Bottlenecks

The speed advantage stems from Cerebras' proprietary Wafer-Scale Engine, a chip that encompasses an entire silicon wafer and was specifically designed to enable faster and more efficient AI model training and inference than traditional GPUs

2

. The architecture keeps model weights on-chip with 44 GB of SRAM on each wafer-sized chip, eliminating the memory-bandwidth bottlenecks that constrain frontier-model inference speed on conventional hardware

1

. This on-chip approach fundamentally differs from GPU-based systems that must shuttle data between separate memory and processing units. Cerebras offers deployment services to assist customers with data preparation, model architecture design, training management, and inference optimization, alongside a subscription-based software updates service providing ongoing enhancements for hardware purchasers

2

.

Strategic Rollout Focuses on Customer Value Discovery

Sachin Katti, VP Compute Strategy & GPT-Infra at OpenAI, indicated the companies are "exploring what becomes possible when customers can get the intelligence of our most capable models with significantly lower latency"

1

. OpenAI plans to begin with a small group of customers to identify where the speed creates the most value before expanding the service more broadly

1

. This measured approach suggests both companies recognize that different use cases will benefit variably from ultra-low latency. Applications like real-time coding assistants, interactive customer support, and live translation services stand to gain immediately, while batch processing tasks may see less dramatic improvements. As organizations increasingly demand AI systems that respond at human conversational speeds, the partnership between Cerebras Systems and OpenAI positions both companies to capture enterprise customers prioritizing performance alongside intelligence. The limited preview phase will likely inform pricing strategies and help identify which industries derive maximum competitive advantage from the technology.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved