5 Sources
[1]
OpenAI introduces 'Ultrafast,' a new mode that makes GPT 5.6 Sol work at 14x the speed
If you've ever found yourself wishing that ChatGPT was a little bit quicker on the uptake, OpenAI seems to be answering your prayers. The AI lab has rolled out a new mode called Ultrafast, which it says is designed to seriously accelerate the pace at which its latest and most powerful model, GPT 5.6 Sol, accomplishes its work. The company says that Ultrafast can work at 14x the speed of standard processing, delivering up to 750 output tokens -- such tokens represent the distinct pieces of text generated by an LLM when it interacts with a human -- per second. "Until now, getting real-time speed typically meant choosing a smaller or more specialized model," the company said in a blog post on Thursday. "Ultrafast points to progress in a new direction: more useful work per second." OpenAI's competitors, like Anthropic, have similarly launched accelerated versions of their models. Claude has fast mode, although it doesn't deliver the kind of speed that OpenAI is offering here. OpenAI suggests that this high-octane version of GPT 5.6 Sol can be deployed across a number of different corporate workflows, most notably incident response, customer service and support, financial market analysis, and e-commerce, among other relevant areas. Ultrafast, which is currently being released in preview, is being powered by OpenAI's partnership with chipmaker Cerebras. Currently, that preview is only being made available to a small group of customers, although OpenAI says that it will expand access to the feature as "capacity grows."
[2]
OpenAI's new Ultrafast mode runs GPT-5.6 Sol 14 times faster, on Cerebras chips
The preview is a bet that latency, not just intelligence, is what will make AI agents genuinely usable, and a marquee win for the wafer-scale chipmaker Cerebras. OpenAI wants its cleverest model to also be its quickest. The company has previewed Ultrafast, a new tier of its API that runs the flagship GPT-5.6 Sol at up to 14 times the usual speed, reaching around 750 output tokens a second, on hardware built by the wafer-scale chipmaker Cerebras. Ultrafast is not a new model so much as a new way to serve an existing one. It leans on Cerebras's outsized chips to strip out the latency that has long dogged frontier AI, and OpenAI opened a limited preview on 13 August to a small group of customers, with plans to widen access as capacity allows. The pitch turns on a trade-off OpenAI says it can finally dissolve. Until now, anyone who wanted genuinely real-time responses had to drop down to a smaller, less capable model, accepting less intelligence in exchange for speed. Ultrafast is meant to deliver frontier-grade reasoning and near-instant answers at once, rather than forcing a choice between them. That combination matters most for the agentic software the whole industry is chasing. An AI agent that has to think for thirty seconds before every step is a demo, whereas one that answers in the time it takes to hold a conversation starts to feel like a product. Speed, in other words, is quietly becoming a feature as important as raw cleverness. OpenAI is aiming the tier squarely at time-sensitive work, including incident response and debugging, financial research and fraud detection, real-time customer support and voice, and e-commerce. Early testers such as Jane Street, Podium, Basis and Rogo describe the change as qualitative rather than incremental, with one saying that speed "completely changes the call experience for complex work" and unlocks "synchronous experiences for users that were previously limited by intelligence." For Cerebras, the deal is a marquee endorsement at a helpful moment. The company went public in one of the year's biggest listings but has since struggled to convince the market that wafer-scale ambition translates into durable profit, so powering OpenAI's fastest tier is precisely the kind of validation it needed. It is also a reminder that the exotic chip architectures once dismissed as science projects are now doing real work for the biggest names in AI. The speed itself comes from an unusual piece of engineering. Cerebras builds processors the size of a dinner plate, cut from a single silicon wafer, which lets an entire model sit on one chip rather than being split across racks of Nvidia GPUs that must constantly shuttle data between them. Stripping out that internal traffic is what collapses the delay between a prompt and a reply, and it is the basis of the company's long-running argument that its design suits inference far better than the general-purpose chips built for training. The move fits a broader shift, too. As raw model capability begins to plateau, the contest is moving toward who can run those models fastest and most cheaply, a race that has lifted inference specialists like Groq and turned latency into a selling point. Rivals such as SambaNova and a clutch of specialist clouds are chasing the same prize, and the market increasingly rewards whoever can make a given model answer soonest, not simply whoever trained the biggest one. For OpenAI, leaning on Cerebras is also a quiet step away from total dependence on Nvidia, of a piece with its work on its own custom silicon. OpenAI has not published pricing, and running its top model at 14 times the speed on specialist hardware is unlikely to come cheap, so Ultrafast may remain a premium option for latency-obsessed cases rather than a default setting. It is also just a preview, gated to a handful of customers while OpenAI hunts for capacity. Even so, the message is plain enough. In the next phase of the AI race, being clever will not count for much if you are also slow.
[3]
Google and OpenAI Debut Super Fast AI Models
The speed race has shifted from raw intelligence to real-time agents, but only Google's model reaches every developer today. Google and OpenAI both pushed the same message today: AI is now fast enough to feel impressive for those who use AI agents, each company announcing ultra fast models. The two launches are built differently, and only one is actually in your hands today. Google shipped Gemini 3.7 Flash, its latest model tuned for software coding and autonomous business workflows. OpenAI opened a limited preview of GPT-5.6 Sol Ultrafast, a new service tier that runs its most capable model at up to 750 output tokens per second. Gemini 3.7 Flash is a general-availability model. It takes up to a million input tokens (roughly 750,000 words) and returns 64,000, handling text, images, video, audio and PDFs, and it can call tools and control a computer. Google is pitching it as the cheap brain for autonomous systems that plan tasks and finish multi-step jobs with less human help. The model is not sacrificing quality for speed. It is both more capable and more efficient, being able to complete our test coding task in 2 minutes and 13 seconds whereas the latest Flash model took more than 5 minutes. The quality gap between the two is also noticeable. OpenAI's Ultrafast isn't a new model. It's GPT-5.6 Sol -- the same model OpenAI used an AI red team to harden against prompt-injection attacks before launch -- on a faster track, powered by chipmaker Cerebras. It is around 14 times faster than GPT-5.6 Sol's own standard speed. Cerebras' wafer-scale chips generate up to 750 tokens a second, about 560 words, fast enough that a voice agent can think mid-call. The numbers that matter Google's own benchmark sheet puts Gemini 3.7 Flash ahead of Claude Sonnet 5, GPT-5.6 Terra, and others on 11 of 18 tested categories, including a top Code Arena web-dev score of 1,588 Elo and 30.4% on AutomationBench for enterprise workflows. That's all based on Google's methodology, so treat the lead as the company's claim. As any Gemini Flash, the model is also cheap. At 75 cents per million input tokens and $3.75 per million output tokens through year-end, it's half of Gemini 3.6 Flash's original rate. That intro price expires December 31, then doubles to $1.50 and $7.50, which is still cheap for a Google model. OpenAI hasn't published head-to-head scores for Ultrafast beyond customer quotes. Jane Street AI engineer John Crepezzi said in OpenAI's announcement that Cerebras' speed "enables different ways of using the models." Podium product lead Courtland Lykins called it "invaluable in our voice stack," saying the speed "completely changes the call experience." The tier is invite-only for now. The speed push lands as the labs pivot from "who's smartest" to "who's fast enough for agents." Google's timing is pointed. Its flagship Gemini 3.5 Pro is still missing, with no release date given, three weeks after 3.6 Flash and days after a DeepMind leadership reshuffle that moved Demis Hassabis aside for deputy Koray Kavukcuoglu. OpenAI, meanwhile, is renting Cerebras' speed rather than waiting on its own stack. Gemini 3.7 Flash is live now in more than 160 countries; GPT-5.6 Sol Ultrafast is still invite-only.
[4]
OpenAI Introduces Ultrafast Tier for GPT-5.6 Sol
* OpenAI has announced an early preview of Ultrafast. * The GPT‑5.6 Sol on Ultrafast mode is powered by Cerebras * It is claimed to generate up to 750 output tokens per second OpenAI has introduced a new Ultrafast mode on Tuesday as a preview. The new tier, powered by 750 output tokens per second, is aimed at reducing the time its GPT-5.6 Sol requires to complete tasks. The ChatGPT maker claims that the Ultrafast mode can run GPT-5.6 Sol significantly faster than standard processing.OpenAI confirmed that it is working with select customers to evaluate the performance of the new tier. GPT-5.6 Sol Gets Ultrafast Mode Through a blog post, OpenAI has announced an early preview of Ultrafast. The AI company states that this new API service tier can run GPT-5.6 Sol up to 14 times faster than standard processing. The GPT‑5.6 Sol on Ultrafast mode is powered by Cerebras and is claimed to generate up to 750 output tokens per second. The company notes that this would allow businesses to build more responsive products, make faster decisions, and bring powerful AI directly into their demanding workflows. "GPT‑5.6 Sol on Ultrafast mode is available in a limited preview today to a select group of customers", said OpenAI. It confirmed the plans to expand access as capacity increases. OpenAI has highlighted potential use cases of the Ultrafast, such as real-time incident response, analysing market signals, assessing transactions, customer support, live research and experimentation. The company confirmed that it is working with an initial group of customers to "understand where this speed makes the biggest difference, and how those learnings can inform our products over time". The company is testing GPT‑5.6 Sol in Ultrafast mode across companies in coding, commerce, financial research, customer support and other interactive applications. The company says the early trials are helping it understand where an order-of-magnitude change in speed creates value in workflows. "We will use these findings to guide deployment as capacity grows", it added. OpenAI developers are testing the GPT‑5.6 Sol in Ultrafast mode internally. It confirmed that the team is using Ultrafast for Incident response, and the new mode is claimed to reduce the delay between observing a signal, testing a hypothesis, and choosing the next action. The internal testing also involves the adoption of Ultrafast for research involving search knowledge sources, query data, and gathering, organising, and summarising information across connected tools.
[5]
Speed Becomes the Product as OpenAI and Google Sell Faster AI | PYMNTS.com
OpenAI's new tier is called Ultrafast, and for now it is a preview open to a small group of customers. It runs the company's GPT-5.6 Sol model up to 14 times faster than the standard tier, at up to 750 output tokens per second, according to a Thursday (Aug. 13) company announcement. The model itself underneath is the same. It just answers faster, running on chips from a company called Cerebras. Early customers include Jane Street, Podium, Basis and Rogo, which are testing it for coding, financial research, customer support, voice and commerce, per the announcement. "Until now, getting real-time speed typically meant choosing a smaller or more specialized model," the announcement said. "Ultrafast points to progress in a new direction: more useful work per second." Google launched its own faster model, Gemini 3.7 Flash, the same day. It costs $0.75 per million input tokens and $3.75 per million output tokens through Dec. 31, half what the previous Flash model cost, then doubles on Jan. 1, 2027, to the price Gemini 3.6 Flash carried all along, according to a Thursday company blog post. Google is also claiming a capability gain, calling 3.7 Flash its "most intelligent workhorse model yet for coding and agents" in the post. Benchmarking firm Artificial Analysis clocked Gemini 3.7 Flash's output at about 340 tokens per second, nearly three times the speed of GPT-5.6 Terra and GLM-5.2, and placed it on the frontier of intelligence versus time per task. Both launches point to the same shift. For two years, AI pricing has mostly come down to how capable a model is and how much a company uses it. Speed is becoming a third feature companies are being asked to pay for on its own. Speed Matters More for Fraud Checks Than for Overnight Reports Not every AI task needs to be fast. A bank checking whether a transaction is fraudulent must decide in a fraction of a second. AI-driven fraud detection has already saved at least $5 million for 42% of card issuers, according to the PYMNTS Intelligence report "Where Payment Decisions Happen: How Issuer Data Is Powering the Next Era of Commerce." A slow fraud check is not just annoying. It can mean letting a fraudulent charge go through before the system catches up. In its announcement, OpenAI named voice applications, customer support, commerce, coding and financial research as the kind of work its fast tier is built for, along with incident response, which covers fraud and security threats that need a fast decision. Compare that to a company running an AI job overnight to sort through old documents. Nobody is waiting on the other end. That job can run on cheaper, slower computers with no real cost to the business. It's the same AI capability, but two different prices, depending only on whether someone is waiting on the answer. AI Pricing Is Splitting Into a Fast Lane and a Slow One This is similar to how companies already buy internet service by paying more for a guaranteed fast connection when it matters, and using a cheaper, slower connection everywhere else. Cerebras, the company powering OpenAI's fast tier, made the same argument in its Thursday announcement, comparing the OpenAI launch to earlier tech shifts like the move from dial-up internet to broadband. If that comparison holds, businesses will start splitting their AI spending into two lanes. Fast, expensive AI will be used for jobs where a delay costs real money, like a customer chatting live with a company or a fraud check happening in real time. Slow, cheaper AI will be used for jobs where nobody notices the wait, like an overnight report or a batch of paperwork. This week's launches from OpenAI and Google suggest both companies expect that split to become normal, something businesses will soon plan for deliberately rather than treat as an afterthought. For all PYMNTS AI coverage, subscribe to the daily AI Newsletter.
Share
Copy Link
OpenAI has launched Ultrafast, a new API service tier that runs GPT-5.6 Sol at 14 times standard speed, reaching 750 output tokens per second on Cerebras chips. The limited preview targets time-sensitive corporate workflows including incident response, customer service, and financial analysis, marking a shift in the AI race from intelligence to speed.
OpenAI has introduced Ultrafast, a new API service tier that accelerates GPT-5.6 Sol to run at 14 times standard processing speed, generating up to 750 output tokens per second
1
4
. The preview launched on August 13 to a select group of customers, with plans to expand access as capacity grows. Powered by Cerebras chips, Ultrafast represents a fundamental shift in how AI companies are positioning their offerings, moving from a singular focus on intelligence to delivering speed as a distinct product feature2
.
Source: The Next Web
The faster AI models achieve their performance through Cerebras' wafer-scale chips, which are built the size of a dinner plate and cut from a single silicon wafer
2
. This architecture allows an entire model to sit on one chip rather than being distributed across multiple Nvidia GPUs that must constantly shuttle data between them. By eliminating this internal traffic, Cerebras collapses the delay between a prompt and reply, enabling the 750 tokens per second output rate that makes real-time applications viable. The partnership represents a strategic move by OpenAI to reduce dependence on Nvidia while validating Cerebras' position in the market following its recent public listing2
.Ultrafast targets time-sensitive corporate workflows where latency optimization matters most, including incident response, customer service, financial analysis, fraud detection, and e-commerce operations
1
4
. Early testers including Jane Street, Podium, Basis, and Rogo describe the speed improvement as qualitative rather than incremental. John Crepezzi from Jane Street noted the speed "enables different ways of using the models," while Podium's Courtland Lykins stated it "completely changes the call experience for complex work" and unlocks "synchronous experiences for users that were previously limited by intelligence"2
3
.On the same day OpenAI announced Ultrafast, Google launched Gemini 3.7 Flash, positioning it as its "most intelligent workhorse model yet for coding and agents"
3
. Unlike OpenAI's limited preview, Gemini 3.7 Flash reached general availability immediately in more than 160 countries. The model processes up to one million input tokens and returns 64,000, handling text, images, video, audio, and PDFs. Google priced it at $0.75 per million input tokens and $3.75 per million output tokens through December 31, half the cost of Gemini 3.6 Flash, before doubling to $1.50 and $7.50 on January 1, 20273
. Benchmarking firm Artificial Analysis measured Gemini 3.7 Flash at approximately 340 tokens per second, nearly three times faster than GPT-5.6 Terra5
.
Source: PYMNTS
Related Stories
The simultaneous launches signal a fundamental restructuring of AI pricing models, where speed becomes a third feature companies pay for independently alongside capability and usage volume
5
. OpenAI acknowledged this shift, stating "until now, getting real-time speed typically meant choosing a smaller or more specialized model" but that "Ultrafast points to progress in a new direction: more useful work per second"1
. The AI model efficiency gains matter differently across use cases. A bank checking transaction fraud must decide in fractions of a second, with AI-driven fraud detection already saving at least $5 million for 42% of card issuers according to PYMNTS Intelligence research5
. Contrast this with overnight document processing where nobody waits for results, suggesting businesses will split AI spending into fast and slow lanes based on whether delays cost real money.The speed race reflects broader market dynamics as raw model capability begins to plateau, shifting competition toward who can run models fastest and most cheaply
2
. Rivals including SambaNova, Groq, and specialist clouds are pursuing similar latency optimization strategies. For agentic AI systems to feel like products rather than demos, they need to respond in conversational time rather than making users wait 30 seconds between steps. OpenAI has not published pricing for Ultrafast, and running its top model at 14 times speed on specialist hardware likely commands premium pricing, positioning it as a high-end option for latency-critical applications rather than a default setting2
. Watch for how quickly OpenAI expands access beyond the initial customer group, whether pricing details emerge that clarify the cost-benefit calculation for businesses, and how competitors respond to this new speed-focused positioning in the AI race.
Source: Decrypt
Summarized by
Navi
[1]
[3]
[4]
12 Feb 2026•Technology

14 Aug 2026•Business and Economy

17 Mar 2026•Technology

1
Policy and Regulation

2
Technology

3
Technology
