5 Sources
[1]
OpenAI introduces 'Ultrafast,' a new mode that makes GPT 5.6 Sol work at 14x the speed
If you've ever found yourself wishing that ChatGPT was a little bit quicker on the uptake, OpenAI seems to be answering your prayers. The AI lab has rolled out a new mode called Ultrafast, which it says is designed to seriously accelerate the pace at which its latest and most powerful model, GPT
[2]
OpenAI's new Ultrafast mode runs GPT-5.6 Sol 14 times faster, on Cerebras chips
The preview is a bet that latency, not just intelligence, is what will make AI agents genuinely usable, and a marquee win for the wafer-scale chipmaker Cerebras. OpenAI wants its cleverest model to also be its quickest. The company has previewed Ultrafast, a new tier of its API that runs the
[3]
Google and OpenAI Debut Super Fast AI Models
The speed race has shifted from raw intelligence to real-time agents, but only Google's model reaches every developer today. Google and OpenAI both pushed the same message today: AI is now fast enough to feel impressive for those who use AI agents, each company announcing ultra fast models. The
[4]
OpenAI Introduces Ultrafast Tier for GPT-5.6 Sol
* OpenAI has announced an early preview of Ultrafast. * The GPT‑5.6 Sol on Ultrafast mode is powered by Cerebras * It is claimed to generate up to 750 output tokens per second OpenAI has introduced a new Ultrafast mode on Tuesday as a preview. The new tier, powered by 750 output tokens per
[5]
Speed Becomes the Product as OpenAI and Google Sell Faster AI | PYMNTS.com
OpenAI's new tier is called Ultrafast, and for now it is a preview open to a small group of customers. It runs the company's GPT-5.6 Sol model up to 14 times faster than the standard tier, at up to 750 output tokens per second, according to a Thursday (Aug. 13) company announcement. The model
Share
Copy Link
OpenAI has launched Ultrafast, a new API service tier that runs GPT-5.6 Sol at 14 times standard speed, reaching 750 output tokens per second on Cerebras chips. The limited preview targets time-sensitive corporate workflows including incident response, customer service, and financial analysis, marking a shift in the AI race from intelligence to speed.
OpenAI has introduced Ultrafast, a new API service tier that accelerates GPT-5.6 Sol to run at 14 times standard processing speed, generating up to 750 output tokens per second
1
4
. The preview launched on August 13 to a select group of customers, with plans to expand access as capacity grows. Powered by Cerebras chips, Ultrafast represents a fundamental shift in how AI companies are positioning their offerings, moving from a singular focus on intelligence to delivering speed as a distinct product feature2
.
Source: The Next Web
The faster AI models achieve their performance through Cerebras' wafer-scale chips, which are built the size of a dinner plate and cut from a single silicon wafer
2
. This architecture allows an entire model to sit on one chip rather than being distributed across multiple Nvidia GPUs that must constantly shuttle data between them. By eliminating this internal traffic, Cerebras collapses the delay between a prompt and reply, enabling the 750 tokens per second output rate that makes real-time applications viable. The partnership represents a strategic move by OpenAI to reduce dependence on Nvidia while validating Cerebras' position in the market following its recent public listing2
.Ultrafast targets time-sensitive corporate workflows where latency optimization matters most, including incident response, customer service, financial analysis, fraud detection, and e-commerce operations
1
4
. Early testers including Jane Street, Podium, Basis, and Rogo describe the speed improvement as qualitative rather than incremental. John Crepezzi from Jane Street noted the speed "enables different ways of using the models," while Podium's Courtland Lykins stated it "completely changes the call experience for complex work" and unlocks "synchronous experiences for users that were previously limited by intelligence"2
3
.On the same day OpenAI announced Ultrafast, Google launched Gemini 3.7 Flash, positioning it as its "most intelligent workhorse model yet for coding and agents"
3
. Unlike OpenAI's limited preview, Gemini 3.7 Flash reached general availability immediately in more than 160 countries. The model processes up to one million input tokens and returns 64,000, handling text, images, video, audio, and PDFs. Google priced it at $0.75 per million input tokens and $3.75 per million output tokens through December 31, half the cost of Gemini 3.6 Flash, before doubling to $1.50 and $7.50 on January 1, 20273
. Benchmarking firm Artificial Analysis measured Gemini 3.7 Flash at approximately 340 tokens per second, nearly three times faster than GPT-5.6 Terra5
.
Source: PYMNTS
Related Stories
The simultaneous launches signal a fundamental restructuring of AI pricing models, where speed becomes a third feature companies pay for independently alongside capability and usage volume
5
. OpenAI acknowledged this shift, stating "until now, getting real-time speed typically meant choosing a smaller or more specialized model" but that "Ultrafast points to progress in a new direction: more useful work per second"1
. The AI model efficiency gains matter differently across use cases. A bank checking transaction fraud must decide in fractions of a second, with AI-driven fraud detection already saving at least $5 million for 42% of card issuers according to PYMNTS Intelligence research5
. Contrast this with overnight document processing where nobody waits for results, suggesting businesses will split AI spending into fast and slow lanes based on whether delays cost real money.The speed race reflects broader market dynamics as raw model capability begins to plateau, shifting competition toward who can run models fastest and most cheaply
2
. Rivals including SambaNova, Groq, and specialist clouds are pursuing similar latency optimization strategies. For agentic AI systems to feel like products rather than demos, they need to respond in conversational time rather than making users wait 30 seconds between steps. OpenAI has not published pricing for Ultrafast, and running its top model at 14 times speed on specialist hardware likely commands premium pricing, positioning it as a high-end option for latency-critical applications rather than a default setting2
. Watch for how quickly OpenAI expands access beyond the initial customer group, whether pricing details emerge that clarify the cost-benefit calculation for businesses, and how competitors respond to this new speed-focused positioning in the AI race.
Source: Decrypt
Summarized by
Navi
[1]
[3]
[4]
12 Feb 2026•Technology

21 Sept 2026•Technology

14 Aug 2026•Business and Economy

1
Technology

2
Technology

3
Policy and Regulation
