5 Sources
[1]
OpenAI voice models get GPT-5-class reasoning
Voice agents have been expensive to run and painful to orchestrate, not because the models can't handle conversation, but because context ceilings forced enterprises to build session resets, state compression, and reconstruction layers into every deployment. OpenAI's three new voice models are
[2]
OpenAI's new voice AI can listen, think, and talk back in 70+ languages
The universal translator just left science fiction and landed in your app store. OpenAI has launched three new audio models in its Realtime API, and they are a big deal for anyone building voice-powered apps. The three models are GPT-Realtime-2, GPT-Realtime-Translate, and
[3]
OpenAI's Brand New Voice AI Is Here. It Could Change How Companies Talk to Their Customers
OpenAI just launched a new set of voice models that can have longer conversations, instantly translate between languages, and more accurately transcribe spoken words into text. The new models are available for businesses to use in their products and services. According to OpenAI, companies
[4]
OpenAI rolls out GPT-Realtime-2, Translate and Whisper audio models for voice AI
OpenAI has introduced three new audio models in its API that enable developers to build a new class of voice applications. These models are designed to make voice interactions more natural, context-aware, and capable of taking action in real time. The three models -- GPT-Realtime-2,
[5]
OpenAI launches 3 advanced realtime voice AI models: Here is what they can do
OpenAI says all three models are now available through its Realtime API. OpenAI has introduced three new realtime voice AI models, which are designed to help developers create smarter and more natural voice-based applications. The new models focus on live conversations, real-time translation, and
Share
Copy Link
OpenAI introduced three specialized voice AI models that bring advanced reasoning to live conversations. GPT-Realtime-2 features GPT-5 class reasoning with a 128K token context window, GPT-Realtime-Translate handles real-time multilingual translation across 70+ languages, and GPT-Realtime-Whisper delivers streaming speech-to-text transcription. Companies like Zillow and Priceline are already building voice agents that can complete complex tasks during ongoing conversations.
OpenAI has launched three new voice AI models that fundamentally change how developers can build voice-powered applications. The company introduced GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper through its Realtime API, separating conversational reasoning, translation, and transcription into specialized components rather than bundling them in a single voice product
1
. This modular approach allows enterprises to assign each task to the appropriate model rather than routing everything through a single, all-encompassing voice system.
Source: Inc.
GPT-Realtime-2 represents OpenAI's first voice model with GPT-5 class reasoning designed specifically for live conversational use cases
1
. The model can handle difficult requests and keep conversations flowing naturally while processing complex agentic tasks. OpenAI expanded the context window from 32K to 128K tokens, enabling longer and more detailed customer interaction without requiring session resets or state compression layers that previously plagued enterprise deployments5
.Developers gain granular control over the model's behavior, including the ability to specify phrases the voice agents should use, adjust reasoning effort based on task complexity, and call multiple tools simultaneously
3
. The model can even narrate its actions with phrases like "checking your calendar" or "let me look into that," making interactions feel more natural and transparent2
.
Source: FoneArena
GPT-Realtime-Translate functions as a universal translator, supporting real-time multilingual translation across more than 70 input languages and 13 output languages
2
. The model translates speech instantly while preserving meaning and pacing, maintaining accuracy even with interruptions, accent variations, or context switching4
. Deutsche Telekom is testing the technology for scenarios where users speak different languages while the system translates conversations instantly with low latency4
.GPT-Realtime-Whisper delivers low-latency speech-to-text transcription by converting spoken audio into text as the speaker talks, unlike traditional models that wait for the speaker to finish
2
. This streaming speech-to-text capability enables real-time understanding for live captions, meeting notes, and voice-powered workflows where immediate transcription is essential5
.Related Stories
Companies including Zillow, Priceline, Deutsche Telekom, Vimeo, and Glean are already building business applications with these real-time voice AI models
3
. Zillow is developing a voice assistant that can search homes and schedule tours from a single spoken request2
. Priceline is working toward full trip management through voice, enabling users to check flights and hotels, cancel them, book new ones, handle delays, and receive TSA updates—all through conversational reasoning4
.
Source: VentureBeat
All three models are now available through OpenAI's Realtime API. GPT-Realtime-2 costs $32 per 1 million audio input tokens and $64 per 1 million audio output tokens
5
. GPT-Realtime-Translate is priced at $0.034 per minute, while GPT-Realtime-Whisper costs $0.017 per minute2
. Developers can test the models in the OpenAI Playground or integrate them into applications immediately4
.Organizations evaluating these models will need to consider their orchestration architecture and whether their stack can route discrete voice tasks to specialized models and manage state across the expanded context window
1
. The new models compete against alternatives like Mistral's Voxtral models, which also separate transcription and target enterprise use cases1
.Summarized by
Navi
[1]
[2]
1
Science and Research

2
Policy and Regulation

3
Technology