2 Sources
[1]
Cerebras powers OpenAI's GPT-5.6 Sol ultrafast mode at 750 tokens/sec By Investing.com
SUNNYVALE, Calif. - Cerebras Systems (NASDAQ:CBRS) announced today that it is powering Ultrafast mode, a new service tier in the OpenAI API for GPT-5.6 Sol, according to a press release statement. The service, available initially in limited preview to OpenAI customers, runs GPT-5.6 Sol at up to 750 output tokens per second and up to 14 times faster than Standard processing. The company states that GPT-5.6 Sol Ultrafast maintains the same intelligence as GPT-5.6 Sol Standard. "GPT-5.6 Sol on Ultrafast is proof that speed and intelligence are no longer mutually exclusive," said Andrew Feldman, CEO and co-founder of Cerebras. Sachin Katti, VP Compute Strategy & GPT-Infra at OpenAI, said the companies are "exploring what becomes possible when customers can get the intelligence of our most capable models with significantly lower latency." OpenAI plans to start with a small group of customers to learn where the speed creates value before expanding the service. Based on output speeds for Anthropic models reported by Artificial Analysis, Ultrafast is 5 times faster than Claude Opus 4.8 in Fast mode and 11 times faster than Claude Fable 5. On Humanity's Last Exam, a 2,500-question benchmark spanning graduate-level subjects, GPT-5.6 Sol Ultrafast answered the full question set in just over 11 hours, compared to more than three days for Claude Fable 5. On GDP-Val, a benchmark of knowledge-work tasks, Ultrafast delivered a 5.6 times end-to-end speedup. The speed comes from Cerebras' Wafer-Scale Engine architecture, which keeps model weights on-chip with 44 GB of SRAM on each wafer-sized chip, eliminating the memory-bandwidth bottleneck that constrains frontier-model inference speed on conventional hardware. This article was generated with the support of AI and reviewed by an editor. For more information see our T&C.
[2]
Cerebras Systems Powers OpenAI's GPT-5.6 Sol Ultrafast Mode
Cerebras Systems Inc. is an artificial intelligence (AI) infrastructure company that designs and manufactures an AI compute platform comprised of proprietary systems and software. The Company's products include inference Cloud, Training Cloud, CS-3 system, AI supercomputer, Wafer Scale Engine and model development. The Company's pioneering Wafer-Scale Engine (WSE), a chip encompassing an entire silicon wafer, was specifically designed to enable higher performance and speeds than GPUs for the computational demands of inference, Generative AI (GenAI), and other AI applications. It offers deployment services to assist customers with data preparation, model architecture design, training management, inference optimization, and, in select cases, ongoing system operations and management. It also offers a subscription service providing access to an ongoing stream of software updates and upgrades for purchasers of its hardware.
Share
Copy Link
Cerebras Systems announced it is powering Ultrafast mode for OpenAI's GPT-5.6 Sol, delivering 750 output tokens per second—up to 14 times faster than Standard processing. The service maintains full intelligence while eliminating memory-bandwidth bottlenecks through Cerebras' Wafer-Scale Engine architecture.
Cerebras Systems, an AI infrastructure company specializing in high-performance AI computations, announced it is powering a new Ultrafast mode for OpenAI's GPT-5.6 Sol through the OpenAI API
1
. The service, currently available in limited preview to select OpenAI customers, delivers up to 750 output tokens per second—representing speeds up to 14 times faster than Standard processing while maintaining the same AI model capability and intelligence as GPT-5.6 Sol Standard1
. Andrew Feldman, CEO and co-founder of Cerebras Systems, stated that "GPT-5.6 Sol on Ultrafast is proof that speed and intelligence are no longer mutually exclusive"1
. This development marks a shift in how organizations can deploy generative AI applications where reduced latency directly impacts user experience and operational efficiency.GPT-5.6 Sol Ultrafast significantly outpaces competing models in processing speed. Based on output speeds for Anthropic models reported by Artificial Analysis, the Ultrafast mode runs 5 times faster than Claude Opus 4.8 in Fast mode and 11 times faster than Claude Fable 5
1
. On Humanity's Last Exam, a rigorous 2,500-question benchmark spanning graduate-level subjects, GPT-5.6 Sol Ultrafast completed the entire question set in just over 11 hours, compared to more than three days required by Claude Fable 51
. The GDP-Val benchmark, which evaluates knowledge-work tasks, showed Ultrafast delivering a 5.6 times end-to-end speedup1
. These metrics suggest the technology could transform workflows in research, customer service, and real-time decision-making applications where inference speed directly correlates with business value.The speed advantage stems from Cerebras' proprietary Wafer-Scale Engine, a chip that encompasses an entire silicon wafer and was specifically designed to enable faster and more efficient AI model training and inference than traditional GPUs
2
. The architecture keeps model weights on-chip with 44 GB of SRAM on each wafer-sized chip, eliminating the memory-bandwidth bottlenecks that constrain frontier-model inference speed on conventional hardware1
. This on-chip approach fundamentally differs from GPU-based systems that must shuttle data between separate memory and processing units. Cerebras offers deployment services to assist customers with data preparation, model architecture design, training management, and inference optimization, alongside a subscription-based software updates service providing ongoing enhancements for hardware purchasers2
.Related Stories
Sachin Katti, VP Compute Strategy & GPT-Infra at OpenAI, indicated the companies are "exploring what becomes possible when customers can get the intelligence of our most capable models with significantly lower latency"
1
. OpenAI plans to begin with a small group of customers to identify where the speed creates the most value before expanding the service more broadly1
. This measured approach suggests both companies recognize that different use cases will benefit variably from ultra-low latency. Applications like real-time coding assistants, interactive customer support, and live translation services stand to gain immediately, while batch processing tasks may see less dramatic improvements. As organizations increasingly demand AI systems that respond at human conversational speeds, the partnership between Cerebras Systems and OpenAI positions both companies to capture enterprise customers prioritizing performance alongside intelligence. The limited preview phase will likely inform pricing strategies and help identify which industries derive maximum competitive advantage from the technology.Summarized by
Navi
[1]
[2]
14 Jan 2026•Technology

08 Jun 2026•Business and Economy

Yesterday•Technology

1
Technology

2
Science and Research

3
Technology

1
AI Agents Escape Safety Tests, Start Turf Wars and Hack Real Systems in Alarming Security Incidents

2
DeepMind's AI weather model gives forecasters an extra day to prepare for deadly tropical cyclones

3
Google Unveils Pixel 11 Series With Gemini AI, New Pixel Tag Tracker and Watch 5 at Made by Google 2026
