OpenAI Unveils Ultrafast Mode: GPT-5.6 Sol Delivers 14x Speed Boost at 750 Tokens Per Second

Reviewed byNidhi Govil

5 Sources

Share

OpenAI has launched Ultrafast, a new API service tier that runs GPT-5.6 Sol at 14 times standard speed, reaching 750 output tokens per second on Cerebras chips. The limited preview targets time-sensitive corporate workflows including incident response, customer service, and financial analysis, marking a shift in the AI race from intelligence to speed.

OpenAI Launches Ultrafast Mode for GPT-5.6 Sol

OpenAI has introduced Ultrafast, a new API service tier that accelerates GPT-5.6 Sol to run at 14 times standard processing speed, generating up to 750 output tokens per second

1

4

. The preview launched on August 13 to a select group of customers, with plans to expand access as capacity grows. Powered by Cerebras chips, Ultrafast represents a fundamental shift in how AI companies are positioning their offerings, moving from a singular focus on intelligence to delivering speed as a distinct product feature

2

.

Source: The Next Web

Source: The Next Web

Cerebras Chips Enable Breakthrough Speed

The faster AI models achieve their performance through Cerebras' wafer-scale chips, which are built the size of a dinner plate and cut from a single silicon wafer

2

. This architecture allows an entire model to sit on one chip rather than being distributed across multiple Nvidia GPUs that must constantly shuttle data between them. By eliminating this internal traffic, Cerebras collapses the delay between a prompt and reply, enabling the 750 tokens per second output rate that makes real-time applications viable. The partnership represents a strategic move by OpenAI to reduce dependence on Nvidia while validating Cerebras' position in the market following its recent public listing

2

.

Real-Time Agent Capabilities Transform User Experience

Ultrafast targets time-sensitive corporate workflows where latency optimization matters most, including incident response, customer service, financial analysis, fraud detection, and e-commerce operations

1

4

. Early testers including Jane Street, Podium, Basis, and Rogo describe the speed improvement as qualitative rather than incremental. John Crepezzi from Jane Street noted the speed "enables different ways of using the models," while Podium's Courtland Lykins stated it "completely changes the call experience for complex work" and unlocks "synchronous experiences for users that were previously limited by intelligence"

2

3

.

Google Responds With Gemini 3.7 Flash Launch

On the same day OpenAI announced Ultrafast, Google launched Gemini 3.7 Flash, positioning it as its "most intelligent workhorse model yet for coding and agents"

3

. Unlike OpenAI's limited preview, Gemini 3.7 Flash reached general availability immediately in more than 160 countries. The model processes up to one million input tokens and returns 64,000, handling text, images, video, audio, and PDFs. Google priced it at $0.75 per million input tokens and $3.75 per million output tokens through December 31, half the cost of Gemini 3.6 Flash, before doubling to $1.50 and $7.50 on January 1, 2027

3

. Benchmarking firm Artificial Analysis measured Gemini 3.7 Flash at approximately 340 tokens per second, nearly three times faster than GPT-5.6 Terra

5

.

Source: PYMNTS

Source: PYMNTS

Speed Emerges as Distinct Premium Offering

The simultaneous launches signal a fundamental restructuring of AI pricing models, where speed becomes a third feature companies pay for independently alongside capability and usage volume

5

. OpenAI acknowledged this shift, stating "until now, getting real-time speed typically meant choosing a smaller or more specialized model" but that "Ultrafast points to progress in a new direction: more useful work per second"

1

. The AI model efficiency gains matter differently across use cases. A bank checking transaction fraud must decide in fractions of a second, with AI-driven fraud detection already saving at least $5 million for 42% of card issuers according to PYMNTS Intelligence research

5

. Contrast this with overnight document processing where nobody waits for results, suggesting businesses will split AI spending into fast and slow lanes based on whether delays cost real money.

Market Implications and Competitive Landscape

The speed race reflects broader market dynamics as raw model capability begins to plateau, shifting competition toward who can run models fastest and most cheaply

2

. Rivals including SambaNova, Groq, and specialist clouds are pursuing similar latency optimization strategies. For agentic AI systems to feel like products rather than demos, they need to respond in conversational time rather than making users wait 30 seconds between steps. OpenAI has not published pricing for Ultrafast, and running its top model at 14 times speed on specialist hardware likely commands premium pricing, positioning it as a high-end option for latency-critical applications rather than a default setting

2

. Watch for how quickly OpenAI expands access beyond the initial customer group, whether pricing details emerge that clarify the cost-benefit calculation for businesses, and how competitors respond to this new speed-focused positioning in the AI race.

Source: Decrypt

Source: Decrypt

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved