OpenAI Launches Ultrafast Mode: GPT 5.6 Sol Now Runs 14x Faster at 750 Tokens Per Second

Reviewed byNidhi Govil

3 Sources

Share

OpenAI unveiled Ultrafast, a new mode that accelerates GPT 5.6 Sol to 14 times standard speed, delivering up to 750 output tokens per second. Powered by partnership with chipmaker Cerebras, the tier targets real-time AI agent interactions in incident response, customer service, and financial analysis. Currently in limited preview with plans to expand access.

OpenAI Accelerates GPT 5.6 Sol with Ultrafast Mode

OpenAI has introduced Ultrafast, a new service tier that runs GPT 5.6 Sol at 14 times standard processing speed, delivering up to 750 output tokens per second

1

2

. The company rolled out this mode on August 13 in limited preview to a small group of customers, with plans to expand access as capacity grows. This represents a shift in how OpenAI positions its flagship model, prioritizing latency optimization alongside intelligence.

Source: The Next Web

Source: The Next Web

Partnership with Chipmaker Cerebras Powers Speed Breakthrough

The Ultrafast mode relies on OpenAI's partnership with chipmaker Cerebras, whose wafer-scale chips enable the dramatic speed increase

1

2

. Cerebras builds processors the size of a dinner plate, cut from a single silicon wafer, allowing an entire model to sit on one chip rather than being split across racks of Nvidia GPUs that must constantly shuttle data between them

2

. Stripping out that internal traffic collapses the delay between a prompt and a reply. For Cerebras, which went public in one of the year's biggest listings, powering OpenAI's fastest tier provides marquee validation at a critical moment

2

.

Real-Time AI Agent Interactions Become Viable

OpenAI positions Ultrafast as solving a longstanding trade-off where real-time speed typically meant choosing a smaller or more specialized model

1

. The company targets the mode at time-sensitive work including incident response and debugging, customer service and support, financial analysis and fraud detection, and e-commerce

1

2

. Early testers including Jane Street, Podium, Basis, and Rogo describe the change as qualitative rather than incremental. Podium product lead Courtland Lykins stated the speed "completely changes the call experience for complex work" and unlocks "synchronous experiences for users that were previously limited by intelligence"

2

3

. Jane Street AI engineer John Crepezzi noted that Cerebras' speed "enables different ways of using the models"

3

.

Source: TechCrunch

Source: TechCrunch

Competition Intensifies Around Super Fast AI Models

The launch coincides with Google shipping Gemini 3.7 Flash, its latest model tuned for software coding and autonomous business workflows, which is available to developers immediately in more than 160 countries

3

. While OpenAI's competitors like Anthropic have launched accelerated versions of their models such as Claude's fast mode, none deliver the speed OpenAI offers with Ultrafast

1

. The speed race has shifted from raw intelligence to real-time applications, with the market increasingly rewarding whoever can make a given model answer soonest

2

3

. Rivals such as SambaNova, Groq, and specialist clouds are chasing the same prize, turning AI model efficiency into a critical selling point

2

.

Strategic Implications for Agentic AI Systems

The move reflects OpenAI's bet that latency, not just intelligence, determines whether agentic AI systems become genuinely usable products

2

. An AI agent that takes thirty seconds before every step remains a demo, whereas one that answers in conversational time starts feeling like a product

2

. For OpenAI, leaning on Cerebras also represents a quiet step away from total dependence on Nvidia, aligning with the company's work on custom silicon

2

. OpenAI has not published pricing for Ultrafast, and running its top model at 14 times the speed on specialist hardware is unlikely to come cheap, suggesting it may remain a premium option for latency-obsessed cases rather than a default setting

2

. As raw model capability begins to plateau, the contest is moving toward who can run those models fastest and most cheaply, with speed becoming as important a feature as raw cleverness

2

3

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved