3 Sources
[1]
OpenAI introduces 'Ultrafast,' a new mode that makes GPT 5.6 Sol work at 14x the speed
If you've ever found yourself wishing that ChatGPT was a little bit quicker on the uptake, OpenAI seems to be answering your prayers. The AI lab has rolled out a new mode called Ultrafast, which it says is designed to seriously accelerate the pace at which its latest and most powerful model, GPT 5.6 Sol, accomplishes its work. The company says that Ultrafast can work at 14x the speed of standard processing, delivering up to 750 output tokens -- such tokens represent the distinct pieces of text generated by an LLM when it interacts with a human -- per second. "Until now, getting real-time speed typically meant choosing a smaller or more specialized model," the company said in a blog post on Thursday. "Ultrafast points to progress in a new direction: more useful work per second." OpenAI's competitors, like Anthropic, have similarly launched accelerated versions of their models. Claude has fast mode, although it doesn't deliver the kind of speed that OpenAI is offering here. OpenAI suggests that this high-octane version of GPT 5.6 Sol can be deployed across a number of different corporate workflows, most notably incident response, customer service and support, financial market analysis, and e-commerce, among other relevant areas. Ultrafast, which is currently being released in preview, is being powered by OpenAI's partnership with chipmaker Cerebras. Currently, that preview is only being made available to a small group of customers, although OpenAI says that it will expand access to the feature as "capacity grows."
[2]
OpenAI's new Ultrafast mode runs GPT-5.6 Sol 14 times faster, on Cerebras chips
The preview is a bet that latency, not just intelligence, is what will make AI agents genuinely usable, and a marquee win for the wafer-scale chipmaker Cerebras. OpenAI wants its cleverest model to also be its quickest. The company has previewed Ultrafast, a new tier of its API that runs the flagship GPT-5.6 Sol at up to 14 times the usual speed, reaching around 750 output tokens a second, on hardware built by the wafer-scale chipmaker Cerebras. Ultrafast is not a new model so much as a new way to serve an existing one. It leans on Cerebras's outsized chips to strip out the latency that has long dogged frontier AI, and OpenAI opened a limited preview on 13 August to a small group of customers, with plans to widen access as capacity allows. The pitch turns on a trade-off OpenAI says it can finally dissolve. Until now, anyone who wanted genuinely real-time responses had to drop down to a smaller, less capable model, accepting less intelligence in exchange for speed. Ultrafast is meant to deliver frontier-grade reasoning and near-instant answers at once, rather than forcing a choice between them. That combination matters most for the agentic software the whole industry is chasing. An AI agent that has to think for thirty seconds before every step is a demo, whereas one that answers in the time it takes to hold a conversation starts to feel like a product. Speed, in other words, is quietly becoming a feature as important as raw cleverness. OpenAI is aiming the tier squarely at time-sensitive work, including incident response and debugging, financial research and fraud detection, real-time customer support and voice, and e-commerce. Early testers such as Jane Street, Podium, Basis and Rogo describe the change as qualitative rather than incremental, with one saying that speed "completely changes the call experience for complex work" and unlocks "synchronous experiences for users that were previously limited by intelligence." For Cerebras, the deal is a marquee endorsement at a helpful moment. The company went public in one of the year's biggest listings but has since struggled to convince the market that wafer-scale ambition translates into durable profit, so powering OpenAI's fastest tier is precisely the kind of validation it needed. It is also a reminder that the exotic chip architectures once dismissed as science projects are now doing real work for the biggest names in AI. The speed itself comes from an unusual piece of engineering. Cerebras builds processors the size of a dinner plate, cut from a single silicon wafer, which lets an entire model sit on one chip rather than being split across racks of Nvidia GPUs that must constantly shuttle data between them. Stripping out that internal traffic is what collapses the delay between a prompt and a reply, and it is the basis of the company's long-running argument that its design suits inference far better than the general-purpose chips built for training. The move fits a broader shift, too. As raw model capability begins to plateau, the contest is moving toward who can run those models fastest and most cheaply, a race that has lifted inference specialists like Groq and turned latency into a selling point. Rivals such as SambaNova and a clutch of specialist clouds are chasing the same prize, and the market increasingly rewards whoever can make a given model answer soonest, not simply whoever trained the biggest one. For OpenAI, leaning on Cerebras is also a quiet step away from total dependence on Nvidia, of a piece with its work on its own custom silicon. OpenAI has not published pricing, and running its top model at 14 times the speed on specialist hardware is unlikely to come cheap, so Ultrafast may remain a premium option for latency-obsessed cases rather than a default setting. It is also just a preview, gated to a handful of customers while OpenAI hunts for capacity. Even so, the message is plain enough. In the next phase of the AI race, being clever will not count for much if you are also slow.
[3]
Google and OpenAI Debut Super Fast AI Models
The speed race has shifted from raw intelligence to real-time agents, but only Google's model reaches every developer today. Google and OpenAI both pushed the same message today: AI is now fast enough to feel impressive for those who use AI agents, each company announcing ultra fast models. The two launches are built differently, and only one is actually in your hands today. Google shipped Gemini 3.7 Flash, its latest model tuned for software coding and autonomous business workflows. OpenAI opened a limited preview of GPT-5.6 Sol Ultrafast, a new service tier that runs its most capable model at up to 750 output tokens per second. Gemini 3.7 Flash is a general-availability model. It takes up to a million input tokens (roughly 750,000 words) and returns 64,000, handling text, images, video, audio and PDFs, and it can call tools and control a computer. Google is pitching it as the cheap brain for autonomous systems that plan tasks and finish multi-step jobs with less human help. The model is not sacrificing quality for speed. It is both more capable and more efficient, being able to complete our test coding task in 2 minutes and 13 seconds whereas the latest Flash model took more than 5 minutes. The quality gap between the two is also noticeable. OpenAI's Ultrafast isn't a new model. It's GPT-5.6 Sol -- the same model OpenAI used an AI red team to harden against prompt-injection attacks before launch -- on a faster track, powered by chipmaker Cerebras. It is around 14 times faster than GPT-5.6 Sol's own standard speed. Cerebras' wafer-scale chips generate up to 750 tokens a second, about 560 words, fast enough that a voice agent can think mid-call. The numbers that matter Google's own benchmark sheet puts Gemini 3.7 Flash ahead of Claude Sonnet 5, GPT-5.6 Terra, and others on 11 of 18 tested categories, including a top Code Arena web-dev score of 1,588 Elo and 30.4% on AutomationBench for enterprise workflows. That's all based on Google's methodology, so treat the lead as the company's claim. As any Gemini Flash, the model is also cheap. At 75 cents per million input tokens and $3.75 per million output tokens through year-end, it's half of Gemini 3.6 Flash's original rate. That intro price expires December 31, then doubles to $1.50 and $7.50, which is still cheap for a Google model. OpenAI hasn't published head-to-head scores for Ultrafast beyond customer quotes. Jane Street AI engineer John Crepezzi said in OpenAI's announcement that Cerebras' speed "enables different ways of using the models." Podium product lead Courtland Lykins called it "invaluable in our voice stack," saying the speed "completely changes the call experience." The tier is invite-only for now. The speed push lands as the labs pivot from "who's smartest" to "who's fast enough for agents." Google's timing is pointed. Its flagship Gemini 3.5 Pro is still missing, with no release date given, three weeks after 3.6 Flash and days after a DeepMind leadership reshuffle that moved Demis Hassabis aside for deputy Koray Kavukcuoglu. OpenAI, meanwhile, is renting Cerebras' speed rather than waiting on its own stack. Gemini 3.7 Flash is live now in more than 160 countries; GPT-5.6 Sol Ultrafast is still invite-only.
Share
Copy Link
OpenAI unveiled Ultrafast, a new mode that accelerates GPT 5.6 Sol to 14 times standard speed, delivering up to 750 output tokens per second. Powered by partnership with chipmaker Cerebras, the tier targets real-time AI agent interactions in incident response, customer service, and financial analysis. Currently in limited preview with plans to expand access.
OpenAI has introduced Ultrafast, a new service tier that runs GPT 5.6 Sol at 14 times standard processing speed, delivering up to 750 output tokens per second
1
2
. The company rolled out this mode on August 13 in limited preview to a small group of customers, with plans to expand access as capacity grows. This represents a shift in how OpenAI positions its flagship model, prioritizing latency optimization alongside intelligence.
Source: The Next Web
The Ultrafast mode relies on OpenAI's partnership with chipmaker Cerebras, whose wafer-scale chips enable the dramatic speed increase
1
2
. Cerebras builds processors the size of a dinner plate, cut from a single silicon wafer, allowing an entire model to sit on one chip rather than being split across racks of Nvidia GPUs that must constantly shuttle data between them2
. Stripping out that internal traffic collapses the delay between a prompt and a reply. For Cerebras, which went public in one of the year's biggest listings, powering OpenAI's fastest tier provides marquee validation at a critical moment2
.OpenAI positions Ultrafast as solving a longstanding trade-off where real-time speed typically meant choosing a smaller or more specialized model
1
. The company targets the mode at time-sensitive work including incident response and debugging, customer service and support, financial analysis and fraud detection, and e-commerce1
2
. Early testers including Jane Street, Podium, Basis, and Rogo describe the change as qualitative rather than incremental. Podium product lead Courtland Lykins stated the speed "completely changes the call experience for complex work" and unlocks "synchronous experiences for users that were previously limited by intelligence"2
3
. Jane Street AI engineer John Crepezzi noted that Cerebras' speed "enables different ways of using the models"3
.
Source: TechCrunch
Related Stories
The launch coincides with Google shipping Gemini 3.7 Flash, its latest model tuned for software coding and autonomous business workflows, which is available to developers immediately in more than 160 countries
3
. While OpenAI's competitors like Anthropic have launched accelerated versions of their models such as Claude's fast mode, none deliver the speed OpenAI offers with Ultrafast1
. The speed race has shifted from raw intelligence to real-time applications, with the market increasingly rewarding whoever can make a given model answer soonest2
3
. Rivals such as SambaNova, Groq, and specialist clouds are chasing the same prize, turning AI model efficiency into a critical selling point2
.The move reflects OpenAI's bet that latency, not just intelligence, determines whether agentic AI systems become genuinely usable products
2
. An AI agent that takes thirty seconds before every step remains a demo, whereas one that answers in conversational time starts feeling like a product2
. For OpenAI, leaning on Cerebras also represents a quiet step away from total dependence on Nvidia, aligning with the company's work on custom silicon2
. OpenAI has not published pricing for Ultrafast, and running its top model at 14 times the speed on specialist hardware is unlikely to come cheap, suggesting it may remain a premium option for latency-obsessed cases rather than a default setting2
. As raw model capability begins to plateau, the contest is moving toward who can run those models fastest and most cheaply, with speed becoming as important a feature as raw cleverness2
3
.Summarized by
Navi
[1]
[3]
12 Feb 2026•Technology

17 Mar 2026•Technology

26 Jun 2026•Policy and Regulation

1
Technology

2
Science and Research

3
Technology

1
AI Agents Escape Safety Tests, Start Turf Wars and Hack Real Systems in Alarming Security Incidents

2
DeepMind's AI weather model gives forecasters an extra day to prepare for deadly tropical cyclones

3
Google Unveils Pixel 11 Series With Gemini AI, New Pixel Tag Tracker and Watch 5 at Made by Google 2026
