29 Sources
[1]
OpenAI releases new voice models for more natural live conversations
OpenAI today released new conversational models, called GPT-Live-1 and GPT-Live-1 mini, claiming that they sound more natural and can handle turn-taking better. These are full-duplex models, meaning they can speak and listen at the same time, allowing users to interrupt naturally and enabling features like live translation. The company is also replacing its current Advanced Voice Mode in ChatGPT with GPT-Live-1 mini by default. Users of paid tiers will be able to access the larger GPT-Live-1 model. The previous model combined a speech-to-text model to transcribe speech, a large language model to generate responses, and a text-to-speech model to deliver the final answer. The company said in a press briefing that the new models solve issues like interrupting users while they're talking and not having enough intelligence to answer questions. OpenAI's new models will send the query to its latest text models like GPT-5.5 for search, reasoning, or agentic capabilities while continuing the conversation. OpenAI also showed that the model can stay silent for a long time and absorb the context of the conversation until it's called upon. Plus, as the new voice mode has access to newer GPT models, it can also present some information in a visual format. Other startups like Monogram, which raised $40 million in seed funding from DST and Lux Capital, are also leaning into visual responses to make assistants more interactive. The company said the new voice mode in ChatGPT is designed to have longer conversations. During the briefing, ChatGPT Voice's product lead, Atty Eleti, said he has had 30- to 40-minute-long conversations with the voice feature during walks. OpenAI thinks that voice could be the primary interface to computing for complex work. Reports have suggested that it could launch a pair of earbuds with AI capabilities this year. However, it didn't provide any information on hardware products. "Over time, we think this will also unlock the ability to use voice as a kind of primary interface to computing, and to manage increasingly complex long-running agentic work. The kind of amazing use cases that we see people using Codex and ChatGPT to accomplish, we think voice can be the future interface to all kinds of work," Eleti said. OpenAI has worked on bolstering voice-based features over the past few years to make ChatGPT's voice mode sound more natural. The company said that more than 150 million people talk to ChatGPT using features like Voice and Dictation. Rivals are also attempting to make assistants more expressive. Both Apple and Amazon have updated their assistants to be more conversational with better context handling. Startups like Sesame, founded by Oculus co-founder Brendan Iribe and Ankit Kumar, also launched AI assistants with more natural conversation while completing tasks in the background. OpenAI is moving in the same direction, aiming to let users talk to its assistant hands-free for a longer time. Despite its claim that the new voice mode sounds more natural, the company emphasized that it's not aiming to make this an AI companion. It noted that the new models have safeguards built in to give age-appropriate responses to teens and provide resources if the conversation turns to topics like self-harm. The new voice mode still needs work. During the demo, when the company showed its live translation feature in Hindi, the assistant had a heavy American accent and spoke in Hindi that was unnatural sounding and had slightly bookish tone. The company said the new mode is optimized for "most spoken languages" but didn't specify which ones.
[2]
ChatGPT's New Voice Models Can 'Listen' and 'Talk' at the Same Time
OpenAI is updating its AI voice technology to match its text-based artificial intelligence. Two new models, GPT-Live-1 and a mini version, are rolling out now to all ChatGPT users. The new model duo relies on OpenAI's latest frontier models. You can select from three different levels of intelligence, depending on how in-depth you want the responses to be. But the biggest update should be that the voice sounds more human and is more conversational, Atty Eleti, product lead for ChatGPT voice, told reporters. One of the biggest changes under the hood is a new process called continuous interaction. This framework allows the AI to simultaneous receive information and produce outputs -- so in the case of voice, it can "listen" and "speak" at the same time. This is a change from previous voice modes that could only respond after you finished speaking. It's helpful if you want to do live, simultaneous translation: You can speak in English and have ChatGPT translate what you're saying into Spanish, Hindi or another language with only a small delay period. Part of the new architecture can also delegate or offload tasks to OpenAI's frontier models. For example, it can answer a question you asked while researching another one. "When GPT-Live has to think hard for a question, it can delegate its reasoning and complex task to GPT-5.5, which can do things in parallel, and this GPT-Live can still remain in conversation with the user," said Kundan Kumar, research lead for the GPT-Live model. When it's done, it "seamlessly weaves" the answer into its response. "This is exactly how humans interact with each other," Eleti added. "We keep the conversation going while we think in the background." It also uses OpenAI's newer AI design tech to create some visual answers, when appropriate, "because sometimes the best answer is displayed, not spoken out loud," Eleti said. Think weather reports, sports scores and other kinds of graphics. These new voice models are available now for all ChatGPT users, including free users and paying subscribers, on your mobile apps and the website. Don't worry if you don't see it right away; OpenAI says it may take a day or two to ensure everyone has access. This is separate from Thursday's expected GPT-5.6 series drop. You'll notice that one way the AI voice tries to mimic human speech is with filler words -- those "ums" "uhs" and "likes" that we use while thinking through our thoughts aloud. Those are new to this update. You can also interrupt it with fewer lags. You can have it sit quietly and listen, only responding when called upon with its name as a wake word. Unlike Siri and Alexa, though, ChatGPT will be actively listening until it's manually disabled. OpenAI automatically opts you out of AI training with voice mode. Audio clips are stored for 30 days -- so the AI has the context from prior conversations -- and can be deleted. Having an AI that feels more human comes with risk, though. We've seen numerous cases (and lawsuits) of the harms of anthropomorphizing AI can wreak, particularly for folks struggling with their mental health. OpenAI says the new models have expanded safeguards and performed better than previous models on "key safety areas" including self-harm, psychosis, violence and sexual content.
[3]
I tested ChatGPT's Live Voice upgrade, and it almost felt human - how to try it
Follow ZDNET: Add us as a preferred source on Google. ZDNET's key takeaways * ChatGPT's Live model aims to enhance voice-based conversations with AI. * With Live, ChatGPT can speak and listen to you at the same time. * ChatGPT can now speak with you almost as if it were a real person. I like having voice conversations with my favorite AIs. Chatting by voice feels more convenient and more engaging than interacting via regular text prompts. But depending on the model, the conversation can still feel stilted. In voice mode, most AIs can tackle only one task at a time -- either speaking to you or listening to you. Try to pause or interrupt, and the conversation can go off the rails. Now, OpenAI has unveiled new voice models for ChatGPT that promise to turn the AI into a more natural and accomplished conversationalist. Also: AI Model Release Tracker: OpenAI's new GPT-Live-1 voice model won't interrupt you Added on Wednesday to the ChatGPT website, the Windows app, and the mobile apps, the new GPT-Live uses a "full-duplex architecture." That simply means it can both speak to you and listen to you at the same time. While you're talking, the AI model shows that it's paying attention by sneaking in phrases like "Yeah" or "Mhmm." Depending on the flow of the conversation, GPT-Live can keep up with you in quick back-and-forth banter or stay silent while it gives you time to collect your thoughts. (Disclosure: Ziff Davis, ZDNET's parent company, filed an April 2025 lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.) While the two of you are chatting away, you can ask ChatGPT to conduct research or carry out a request. The online search is handed off to another model behind the scenes, so the AI can focus on your conversation without interruption. That reminds me of myself when I'm speaking with a friend or relative on the phone who needs technical help, and I'm searching for the topic on the web, all at the same time. The new full-duplex conversational model is available for all ChatGPT users. However, there are two different models depending on whether you're a paid subscriber or a free user. GPT‑Live‑1 is the default model for ChatGPT Voice for Go, Plus, and Pro users. GPT‑Live‑1 mini is the default for free users. Between the two, GPT‑Live‑1 offers higher quality, while the mini model uses fewer resources. Want to take it for a spin? You can try Live Voice on the ChatGPT website, the ChatGPT Windows app, and the iOS and Android mobile apps. The ChatGPT Mac app no longer supports Voice mode, so Mac users will have to turn to the website. The new Live model should automatically be accessible. To check at the site or one of the apps, go to Settings and select Voice. The model should show Live as the new default. Select the drop-down menu, and you can always go back to Advanced or Standard, but I recommend keeping it at Live. Also: How to audit what ChatGPT knows about you - and reclaim your data privacy Another option called Intelligence determines the level of reasoning ChatGPT uses in your voice chats. As the default, Instant is fine for general chats and quick replies. Medium digs deeper into your question or request and takes longer to respond. High is for more complex problem solving and takes the longest to finish. You'll want to keep it set at Instant for the most part unless you're conducting complex research. You can also change the voice based on gender, accent, and other attributes. Since I love anything British, my favorite is Vale with her friendly yet bright British accent. To kick off a voice conversation, select the voice icon to the right of the prompt. After you're done, ChatGPT displays a transcription of your entire chat. Though GPT-Live is technically a smaller upgrade than a brand-new general model, it still promises to enhance the experience of speaking with AI assistants. Does it live up to that promise? Here's what I found. Search the web during a conversation In one conversation, I told ChatGPT that I couldn't find a way to change the aspect ratio in the Camera app on an iPad mini. At first, the AI gave me instructions that worked only on an iPhone. I interrupted it to explain that the suggested steps apply to an iPhone but don't seem to work on an iPad mini. In response, ChatGPT searched the web during our conversation to confirm that the iPad mini doesn't let you change the ratio in the camera app. Also: I connected ChatGPT to my bank, and it's my go-to finance app now - here's how (and why) The AI then gave me instructions for changing the ratio in the Photos app. Again, I interrupted and asked it to find third-party camera apps for the iPad that would let me change the ratio before taking a photo. ChatGPT searched the web while it resumed our conversation about changing this in the Photos app. I then asked it to tell me about the third-party apps it found. Throughout the conversation, ChatGPT handled all interruptions and requests while remaining focused on our chat. Change up a story In another conversation, I asked it to tell me a story about my cat, Mr. Giggles, traveling to the moon and landing there. During the chat, I interrupted it a few times, telling it to slow down, speed up, and even change the story so that Mr. Giggles lands on Mars instead. Each time, the AI easily let me interrupt it and adjusted the story based on my request. Veer off in different directions In another conversation, I told ChatGPT I wanted to discuss classic Hollywood films of the '30s, '40s, and '50s. After the conversation kicked off, the AI mentioned screwball comedies of the '30s, noir films of the '40s, and Technicolor musicals of the '50s and asked me which mood struck me. After I suggested screwball comedies, the conversation veered off in that direction. Along the way, ChatGPT suggested several films in that genre. We explored one particular film that I've seen but don't enjoy as much as others. We explored why that could be, and the AI came up with a great explanation. Beyond just taking the chat in different directions, I thoroughly enjoyed the conversation -- it felt like talking to a fellow film lover. Translate a conversation Finally, I asked ChatGPT to translate a live conversation, specifically one between me speaking English and someone else speaking French. The back-and-forth live translations were quick and fluid, fitting right into the conversation. The next time I need a translator when I'm in another country, I'll definitely give ChatGPT the job. Also: How to use ChatGPT: A beginner's guide to mastering OpenAI's chatbot Any problems? What happens if the TV is on in the background or other people are speaking near you? Could that interrupt your conversation? Two ZDNET editors who tried GPT Live said that their conversations were interrupted by background audio and by another person speaking, forcing them to tell the AI to continue the chat. I tried to replicate that problem but I couldn't. Even with loud TV dialogue playing and someone speaking in the background, none of my conversations were ever interrupted. If you do run into this dilemma, there's not much you can do other than turn down the background audio or try to shush the other person speaking. Final thoughts Otherwise, I was quite impressed with ChatGPT's new Live modes and enjoyed speaking with the AI without the usual impediments. The conversations certainly flowed and felt almost like speaking with a real person. From now on, I'll choose ChatGPT when I want to strike up a conversation with an AI.
[4]
ChatGPT's New Voice Models Aim for More Human-Like Conversations
GPT-Live-1 replaces Advanced Voice Mode as the default voice option in ChatGPT, and can understand long pauses in your speech. OpenAI is updating ChatGPT's Voice Mode with new models that aim to make interacting with AI feel as natural as talking to another human. The new models, GPT-Live-1 and GPT-Live-1-mini, are part of the GPT-Live family. They are built on a full-duplex architecture, which allows them to speak and listen simultaneously, the company says in a blog post. GPT-Live-1 replaces Advanced Voice Mode as the default voice option in ChatGPT. It can understand long pauses in your speech and even add phrases like "mhmm," "yeah," or "got it" to show it's paying attention. Previously, Voice Mode followed a turn-based approach, which often meant that ChatGPT would start responding the moment you paused to think. According to OpenAI, GPT-Live-1 is the smartest voice model it has released to date. When a query requires research, reasoning, or agentic capabilities, it delegates the task to the frontier GPT-5.5 model while continuing to talk to you. "This allows it to keep the conversation going, even as it handles multiple tasks in the background," OpenAI says. As users, you can choose the level of smartness you need from GPT-Live. During interactions, tap Settings on the top right, select the Intelligence button, and choose "Instant for fast responses, or Medium and High when you want ChatGPT to spend more time thinking." For topics like weather, stocks, and sports, you may also see rich visual cards appear in the chat window. The GPT-Live models are currently rolling out to ChatGPT users worldwide on iOS, Android, and web. For Go, Plus, and Pro users, GPT-Live-1 will be the default voice model, while GPT‑Live‑1 mini will be the default for free users. The move doesn't phase out legacy models like Standard and Advanced Voice Mode. You can still access them from the app, OpenAI says.
[5]
OpenAI makes ChatGPT better at banter
OpenAI has released a new voice model that can produce human-sounding speech, or scour the web in response to spoken queries. GPT-Live, according to the company, makes chatbot banter feel more like a real conversation, something of a bold move for a company battling multiple lawsuits alleging mental health harms because people took ChatGPT too seriously. "During conversations, GPT‑Live can show it's paying attention with phrases like 'mhmm' or 'yeah', engage in quick back-and-forth, or just stay quiet when you need a moment to think," the company said in a blog post. "The result is a voice experience that is refreshingly easy to talk to." The company has published a video demonstrating this full duplex experience. It features three women of an age seldom seen at companies like OpenAI but often impacted by the kinds of scams AI technology enables. OpenAI insists that it has expanded its safety testing regime to better assess native audio interactions. And it has published a system card to document its approach. While OpenAI notes that it has policies and protections against voice cloning and impersonation, the company has not disavowed replicating a competing product from former CTO Mira Murati. Murati's company Thinking Machines in May talked up "interaction models" and how they can speak, listen, and search the web at the same time. Two months later, OpenAI has a similar offering. If Apple were involved, we'd say Thinking Machines had been "Sherlocked," a term from a time when copying a startup's product stirred indignation. We'd suggest "Altmanned" as an alternative if it weren't for the global shrug of indifference to frontier model companies capturing the world's intellectual output, laundering it, and reselling it. GPT-Live will delegate queries that require web search to a background model (GPT-5.5 presently) that processes the request while maintaining conversational flow with the user. The company's hope is that this will allow voice interaction to drive more complicated, lengthy agentic workflows - which tend to inflate token usage and billing. Whether an original idea, a parallel innovation, or a sincerely flattering imitation, GPT-Live's full-duplex implementation represents an improvement in model architecture. "Instead of processing a sequence of separate messages, GPT‑Live continuously processes input while generating output," OpenAI explains. "The model can therefore make interaction decisions many times per second: whether to speak, continue listening, pause, interrupt, or invoke a tool." It will be interesting to see whether security researchers find that this approach, of continuously processing input, allows for novel attack opportunities. ChatGPT users can invoke GPT-Live by tapping the "Voice" button. OpenAI contends this experience will result in more natural conversations, better answers, improved listening, and visual feedback. GPT-Live will appear in the iOS and Android ChatGPT apps, and on the web. A more capable version, GPT-Live-1, is the default for ChatGPT Voice for Go, Plus, and Pro users. Free-tier customers have to settle for GPT-Live-1 mini. GPT-Live comes with a caveat - it has been optimized for popular languages and may not work all that well for "certain languages" yet. But now that the model has been made more responsive, any missteps should be noticeable a few milliseconds sooner. ®
[6]
OpenAI launches GPT-Live voice models that listen and speak simultaneously
July 8 (Reuters) - OpenAI on Wednesday launched GPT-Live, a new family of voice models capable of listening and speaking simultaneously in real time. The IPO-bound AI startup said it will roll out two versions of GPT-Live - GPT-Live-1 and GPT-Live-1 mini - to users globally on Wednesday. In May, OpenAI introduced three audio models for its developer platform, aiming to make voice-based software agents more conversational and capable of completing tasks in real time. Reporting by Juby Babu in Mexico City; Editing by Vijay Kishore Our Standards: The Thomson Reuters Trust Principles., opens new tab
[7]
The Future Is Always Listening: OpenAI Says Its New Voice Assistant Is 'One Step Closer to a Truly Accessible AGI'
OpenAI just dropped a new AI model that's designed to sound like an actual, human conversation partner -- and the company says it marks a step closer to AGI. For years, OpenAI has been investing in the development of AI tools that can speak in humanlike voices. That's sometimes led to controversy, as when the company debuted a voice assistant that sounded suspiciously similar to Scarlett Johansson. The various shortcomings of AI assistants, meanwhile, such as their difficulty in understanding context, has become a fertile source of mockery on social media. Now the company is hoping to turn the page on its AI voice assistant efforts -- and, it hopes, bring the technology to a more mainstream audience -- with the newly released GPT-Live-1, which it calls its "smartest voice model yet." Unveiled on Wednesday, the model is able to communicate in lifelike voices, sprinkled with subtle intricacies that inflect actual human speech, like sudden bursts of laughter and short intakes of breath before a sentence. But more importantly for OpenAI's goal of boosting AI voice assistants' appeal, GPT-Live-1 also knows when to keep quiet: It can be a passive fly on the wall, attentively following a conversation without chiming in every ten seconds. During these quiet periods it will sprinkle in the occasional "Right" or "Mmhmm," to remind users that it's listening. GPT-Live-1 defers to GPT-5.5 for more complex requests, theoretically making it useful not just for idle conversations but also for tasks that require a deeper level of reasoning. It also comes with upgraded real-time web search and translation capabilities, which OpenAI is likewise underscoring in the hopes of boosting its appeal as a general-purpose assistant. In short, OpenAI is hoping all these new features will add to users' feeling that they're interacting with something more than just a glorified search engine. "You can even forget you're talking to an AI," Yuchen Zhang, a research engineer at OpenAI, said in Chinese during a livestreamed demo on Wednesday, as GPT-Live-1 translated his words to English. "How should I put it? It's a really, really amazing feeling." The company has gone to such great lengths to underscore the naturalness of GPT-Live-1's conversational abilities that one almost gets the sense the company is positioning it as more of a companion than a traditional, question-answering chatbot. In a marketing video posted to its official X account on Wednesday, OpenAI showed three elderly women using the voice model for a variety of tasks, like getting knitting tips, checking for public transit delays, and translating English to French. The decision to include a cast of elderly people rather than, say, hip-looking 20s-somethings could hint at the company's efforts to position voice assistants as being at least as much of a solution to loneliness as they are a quick means of getting a weekly weather forecast. For the most part, though, OpenAI is pushing its new voice model towards a general audience, promoting it as a more useful and intuitive version of ChatGPT. "This model is one step closer to a truly accessible AGI, a world where talking to AI actually starts to feel like a real conversation," Kundan Kumar, another OpenAI researcher, said during the livestream. It's worth remembering that AGI -- short for artificial general intelligence -- is a term that's thrown around quite loosely these days as both a technical benchmark and a marketing buzzword. The industry lacks a clear, single definition for what true AGI would look like, although there's general agreement that it would need to perform a wide variety of intellectual tasks at least as well as the typical human brain. By dropping it into its livestreamed demo, OpenAI could be trying to plant the idea the public's mind that "true AGI," whatever that means, depends on an ability to speak in humanlike voices. OpenAI said in its announcement blog post that the new model has been built with safeguards to prevent some of the harms caused by earlier AI voice assistants. It will not imitate voices of actual people, for example, and if it detects user voice prompts on dangerous subjects like self-harm or other kinds of violence, it will respond by surfacing information about health and safety resources and, in some cases, end the conversation.
[8]
OpenAI's GPT-Live: ChatGPT voice that listens and talks
OpenAI has rebuilt ChatGPT's voice, and the new GPT-Live models can listen and speak at the same time. They roll out to everyone from today, free users included, with live translation and a more human feel. OpenAI wants you to talk to ChatGPT, not type at it. On 8 July it launched GPT-Live, a new generation of voice models. The company says they make talking to AI feel much closer to a real conversation. Two versions, GPT-Live-1 and a smaller GPT-Live-1 mini, reach ChatGPT users worldwide from today. The headline change sits in the architecture. GPT-Live runs "full-duplex," so it can listen and speak at once. It can drop in an "mhmm" or a "got it," trade quick back-and-forth, or wait while you think. Older voice modes had to wait for you to stop talking first. That often left them cutting in at the wrong moment. Think in the background, keep talking The second shift is delegation. When a question needs a web search or harder reasoning, GPT-Live hands it to a stronger model in the background. That model is GPT-5.5 at launch. GPT-Live keeps chatting while it waits, then folds the answer back in. "This is exactly how humans interact with each other," ChatGPT voice product lead Atty Eleti told reporters, according to CNET. "We keep the conversation going while we think in the background." One demo drew attention. In press briefings, the model handled real-time, simultaneous translation, speaking a running translation as the presenter talked. TechRadar, shown the feature, called it genuinely useful. The model also answers to a wake word, with OpenAI staff using "Hey Chat" in the demo. Smarter, and free for everyone OpenAI calls GPT-Live its smartest voice model yet. Users can pick a reasoning level: Instant, Medium or High. Voice can now surface visual cards for things like weather, stocks and sports. The company claims GPT-Live-1 and the mini beat the old Advanced Voice Mode in five to 10-minute test conversations. It also reports big gains on benchmarks for scientific reasoning and agentic web search. More than 150 million people already use ChatGPT's Voice and Dictation each week, OpenAI says. GPT-Live-1 becomes the default for Go, Plus and Pro subscribers, while free users get the mini. It reaches iOS, Android and the web, though the rollout may take a few days. The human-sounding problem Sounding human cuts both ways. OpenAI says it added voice-specific safety training and real-time safeguards. Those can steer a reply, surface crisis resources, or end a conversation in higher-risk cases, with extra protections for teens. It also stresses that GPT-Live uses preset voices and will not imitate a real person. That caution matters. A model that sounds like a friend earns trust more easily, and misleads more easily too. Critics such as Signal's Meredith Whittaker have warned that chatbots "are not your friends." The field stays crowded as well. Rivals like ElevenLabs and a wave of voice-agent startups chase the same prize. The broader assistant race pushed Amazon to rebuild Alexa. Why it matters This lands the same day OpenAI pressed ahead with a broad GPT-5.6 rollout, part of a fast release cadence. For years, voice served as a convenient extra, not the main way people use ChatGPT. With GPT-Live, OpenAI is betting that talking, not typing, becomes the default. Whether people want an AI that never quite stops listening is the open question.
[9]
ChatGPT has new voice models that make conversations feel much more natural
ChatGPT just became a much closer rival to Google's Gemini Live. OpenAI has released a major ChatGPT Voice update that brings far more natural conversations with AI, not to mention better answers. The upgrade uses new GPT-Live-1 and GPT-Live-1-mini models that revolve around continuous two-way interaction. You can talk over it, and it should be much better at knowing when to respond quickly, acknowledge you with an "mmmhmm," or give you a moment to think. Previous ChatGPT Voice approaches relied on you and the AI taking turns, and in the original version needed three models just to work. This not only led to stiff-sounding back-and-forth chats, but could lead to GPT either losing information or interrupting at the wrong moment. GPT-Live can also delegate intensive tasks like agentic work to other models, such as GPT-5.6 or the imminently public GPT-5.6. The conversation will keep going even when it's processing a complex request, OpenAI says. This makes it easier to use the company's latest frontier models, providing faster and more accurate responses. Like Gemini and Apple's upcoming Siri AI, you'll get cards for visual information, such as maps, weather, or the latest World Cup scores. OpenAI says it's updating its safety guardrails to match. GPT-Live can steer models toward "safer" responses, including teen-appropriate answers and help for potential self-harm. The tech giant also promises to gauge feedback over longer periods, helping it to spot possible issues with "emotionally sensitive" conversations. It won't be allowed to imitate real people, either. What is the GPT-Live release date? Paid users get the best experience GPT-Live is deploying now for ChatGPT Voice users on Android, iOS, and the web. It's not initially available in the desktop app, Codex, temporary chats, or custom GPT implementations. The quality of the model depends on what you're willing to pay. Free users will have conversations using the GPT-Live-1-mini model, while Go, Plus, and Pro users will get the full GPT-Live 1 experience. ChatGPT Business, Enterprise, and Edu users will have to wait until sometime after launch. Screen and video sharing aren't available in any ChatGPT Voice system on launch, OpenAI says. If you need them, you can revert to the Standard and Advanced Voice Mode versions until GPT-Live supports the functionality. If that's not an issue, though, you can finally have a Gemini Live-style chat outside of Google's ecosystem. ChatGPT+ What's included? Unlimited conversations, faster response speed, priority access, and more Brand ChatGPT ChatGPT's AI-supported assistance gets even better with a paid subscription; it Plus tier offers enhanced features including unlimited conversations, faster response speed, priority access, and more. Try for Free Expand Collapse
[10]
ChatGPT might actually be worth talking to now
GPT-Live is available to all ChatGPT users including free accounts, promising more natural and fluid AI voice interactions. AI voice modes are so bad that I typically don't bother. Up to now, the voice modes for ChatGPT, Claude, and Gemini have had to take turns as they listen and talk, making for herky-jerky conversations with lengthy diatribes punctuated by abrupt interruptions. But we have seen some AI startups offering more natural-sounding voice modes, including some that actually listen while they talk, and now OpenAI is debuting an upgraded chat mode that's also designed to be a much better listener. Rolling out now to ChatGPT users, GPT-Live (which I haven't had the chance to try yet) boasts a key feature: a full-duplex architecture, which allows the model to both listen and talk at the same time. Previous ChatGPT voice modes were turn-based, similar to old-school CB radios. That means when I talk, you listen -- "Hey, good buddy, what's your 20? Over" -- and when you talk, I listen -- "I'm at Dunkin Donuts on I-5, over." With GPT-Live (which is coming out in two versions, GPT-Live-1 and GPT-Live-1 mini), ChatGPT now employs a full-duplex configuration like on a standard phone call, meaning I can (if you don't mind too much) interrupt or even talk over you. And because it "continuously processes input while generating output," GPT-Live is better at deciding when to talk, interrupt, or -- crucially -- just shut up and listen. Aside from being able to talk and listen at the same time, GPT-Live can also deploy agents to conduct web searches while it's talking, thus eliminating the "searching the web" pauses on the old ChatGPT voice mode when it was looking up something. OpenAI's GPT-Live models aren't the first we've seen (and heard) with full-duplex capabilities. AI voice startup Thinking Machines has demoed voice "interaction" models that can also listen while they talk. There's also Sesame AI, which has a voice mode capable of deploying search agents while it's still talking. But for most of us, GPT-Live will be the first time we actually give full-duplex AI voice chats a try, and who knows? Maybe ChatGPT will actually be worth talking to now. GPT-Live is rolling out to all ChatGPT users starting today, including free users. The new voice mode hasn't landed on my ChatGPT account yet, but I'll share my first impressions once it does.
[11]
OpenAI Introduces GPT-Live to Make ChatGPT Voice Feel Like a Real Conversation
OpenAI today introduced GPT-Live, which it describes as a new generation of voice models meant to make talking to AI feel more like having a conversation with a real person. GPT-Live is meant to replace the existing ChatGPT voice experience. GPT-Live is able to listen and speak at the same time, and it can show it is paying attention with acknowledgment phrases like "mhmm." The model was built for continuous interaction, and it can make decisions on whether to speak, continue listening, pause, interrupt, or use a tool multiple times per second. Talking with ChatGPT should now feel much more like a real conversation. You can interrupt with a question, pause to gather your thoughts, or ask ChatGPT to slow down. It naturally acknowledges what you're saying with phrases like "mhmm" or "got it," so you know it's following along. We've also remastered the nine distinct voices in ChatGPT for GPT-Live. OpenAI says GPT-Live is its smartest voice model to date, using the latest frontier model (currently GPT-5.5) for web search, deep reasoning, and complex work. While GPT-Live works on a task, it is able to continue a conversation, and then give the results of a task when it's finished. It also works for live translation, and displays rich visual cards for weather, stocks, sports, and more. OpenAI is rolling out GPT-Live-1 and GPT-Live-1 mini to ChatGPT users worldwide starting today. GPT-Live-1 is the default for Go, Plus, and Pro users, while GPT-Live-1 mini is the default for Free users. ChatGPT users can tap the Voice button to talk with ChatGPT and experience GPT-Live. GPT-Live does not yet support voice with video or screen sharing in ChatGPT, but OpenAI is working to add that feature soon.
[12]
Got ChatGPT's new voice mode? Here's how to check -- and 5 things you should try first
* GPT-Live is rolling out to all ChatGPT users now * It can both talk and listen at the same time for more natural chats * Real-time translations are also now possible ChatGPT has a shiny new AI voice model called GPT-Live, which has a number of helpful tricks -- including being able to listen and talk at the same time. It's rolling out to all ChatGPT users now, though OpenAI has acknowledged a number of early bugs. While free users and users on a paid plan do get slightly different models -- GPT-Live-1 mini and GPT-Live-1 respectively -- the updated model should now be appearing in all ChatGPT accounts, with the new features outlined below. The biggest giveaway that you've got the upgrade will be the Live label at the top of voice chats on mobile, and behind the ChatGPT drop-down on the web. Tap or click on these labels and you can still go back to the old voice models, for the time being. There's another way to check the GPT-Live voice model has arrived in your account: on mobile, tap the menu button (top left), then the settings cog (Android) or your profile avatar (iOS), and Voice > Model. On the web, click your profile avatar (bottom left), then Settings > Voice > Model. What to try first The biggest upgrade here is the 'duplex' functionality, so try that first: you can keep talking even after ChatGPT has started answering you, and it should keep up. Second, try interrupting it mid-flow, and it'll adapt its response accordingly. We're almost at the level of the 2013 Spike Jonze movie Her at this stage. Third, ask ChatGPT in voice mode to translate something into a foreign language as you say it out loud. You can then speak out sentences in English, and ChatGPT will do a real time translation for you without hesitating. It's not particularly useful for language learning, but it does show off the capabilities of GPT-Live. Fourth, change the voice and intelligence used -- you can do this via the sliders icon at the top right of voice chats. The voice options are actually the same as they were before, but you can choose between Instant, Medium, and High as the intelligence level. Use Instant for the fastest answers, High for the best answers, and Medium for a compromise. The final thing you can try once you've got the update is to ask questions with visual answers. OpenAI has added a bunch of visual cards to voice mode now, so you get graphics on screen about sports scores, weather forecasts, and places that can be found on a map, for example. Early voice bugs I've been testing out GPT-Live voice mode for a few hours and can report that everything works as advertised. It is, more than ever, like talking to a real person -- right down to the hesitations and the variety in speech patterns. I did experience one or two glitches, but they were few and far between. Over on Reddit, OpenAI's Atty Eleti is answering questions about GPT-Live. One of the main bugs that users seem to be experiencing is related to ChatGPT's memory, which appears to be off limits to voice mode in some cases -- this is an issue that OpenAI is tracking and "actively investigating", and you can find updates on it here. Problems are also being reported when it comes to foreign languages being pronounced in an English accent. Again, this is an issue that's been acknowledged, and which should improve over time according to Eleti. Overall though, the rollout seems to be going relatively smoothly -- and I haven't seen any issues with memory or with accents so far. I'm not sure it's going to make me want to use voice mode any more than I already do (which isn't much), but for heavy voice users it's definitely a big step forward. Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.
[13]
OpenAI launches GPT-Live, a full-duplex voice upgrade that lets ChatGPT talk more like a person
OpenAI on Wednesday launched GPT-Live, a pair of new voice models that fundamentally redesign how people talk to ChatGPT -- replacing the company's existing Advanced Voice Mode with an architecture that can listen and speak simultaneously, much like an actual human conversation. The two models, GPT-Live-1 and GPT-Live-1 mini, are rolling out globally starting today across iOS, Android, and ChatGPT.com. GPT-Live-1 becomes the default voice model for paid ChatGPT users on the Go, Plus, and Pro tiers, while GPT-Live-1 mini serves free-tier users. OpenAI also plans to bring the models to the API, and developers can sign up to be notified. The release marks the third generation of ChatGPT's voice technology in roughly two years -- and OpenAI's clearest bid yet to turn its chatbot into something that feels less like querying a search engine and more like talking to a colleague. Why full-duplex voice changes everything about talking to AI The defining technical advance in GPT-Live is what OpenAI calls a "full-duplex architecture." In telecommunications, full-duplex means both parties on a phone call can talk and listen at the same time. Applied to AI, it means the model continuously processes your incoming audio even while it generates its own spoken response -- no more waiting for a clean silence gap to figure out when you've finished a thought. "Instead of processing a sequence of separate messages, GPT-Live continuously processes input while generating output," OpenAI wrote in its research blog. "The model can therefore make interaction decisions many times per second: whether to speak, continue listening, pause, interrupt, or invoke a tool." In practice, that translates to a voice assistant that can insert conversational acknowledgments -- "mhmm," "yeah," "got it" -- while you're still talking, pick up on a natural pause without jumping in prematurely, and handle rapid interruptions without derailing the entire exchange. OpenAI's previous Advanced Voice Mode, launched to paid users in September 2024, processed and generated audio within a single model but still operated on rigid turn-by-turn exchanges. As OpenAI acknowledged in the announcement, "because turn detection is based on silence, even a brief pause or background noise could be mistaken for the end of turn -- causing the model to interrupt at unnatural times." That brittleness created a product that, while impressive in demos, could be deeply frustrating in extended real-world use. Background chatter in a coffee shop could trigger a response. A thinking pause might get swallowed. The experience felt, as one researcher put it on X shortly after the announcement, like "walkie-talkie turn taking." GPT-Live is designed to end that era. How OpenAI split voice and intelligence into two separate layers GPT-Live introduces a second structural change that may prove just as consequential for enterprise adoption: it decouples the voice interaction layer from the reasoning layer. When a user asks a straightforward question, GPT-Live handles it directly. But when the query demands web search, deeper reasoning, or more complex agentic work, GPT-Live delegates the task to a frontier model running in the background -- at launch, GPT-5.5, the large language model OpenAI released in April -- and continues talking with the user while the computation happens asynchronously. "While it works, GPT-Live can keep talking with you and maintain the flow of conversation," OpenAI explains. "As we release new frontier models, we'll continuously update the model used by GPT-Live." This delegation model is a meaningful architectural bet. Rather than building a single monolithic voice model that tries to be both conversationally fluid and deeply intelligent, OpenAI has split the problem in two: a voice-native model optimized for real-time interaction, and a separate reasoning engine that can be swapped out as the state of the art improves. It is, in effect, a modular design -- one that allows OpenAI to upgrade the intelligence of its voice assistant without retraining the voice model itself. The implications for enterprise and developer workflows are significant. A voice agent built on this architecture could maintain a natural conversation with a customer while simultaneously querying databases, searching the web, or performing multi-step reasoning -- tasks that would have introduced several seconds of dead air under the old pipeline. The three generations of ChatGPT voice, from clunky pipeline to continuous stream To understand how far voice AI has come, it helps to trace the three generations that led to GPT-Live. The original ChatGPT Voice, launched in 2023, used a cascaded pipeline -- a speech-to-text model (Whisper) transcribed what you said, a large language model (GPT-4) generated a text response, and a text-to-speech model converted that response back into audio. Each handoff introduced latency and lost information. As OpenAI noted, "the complexity came at a cost: information could be lost across models, and responses were slow and stilted." That cascaded approach was the industry standard, and its limitations were well-documented. As the blog OpenHelm noted in an October 2024 analysis of OpenAI's Realtime API, the old pipeline stacked up to roughly 1,700 milliseconds of latency -- nearly two full seconds of dead air before the first word of a response. Managing the state between the three separate APIs consumed an enormous amount of engineering effort. OpenAI's Advanced Voice Mode, which began its limited rollout to paid ChatGPT Plus users in July 2024 before expanding more broadly in September 2024, collapsed that three-model pipeline into a single model that processed audio natively. As TechCrunch reported at the time, the rollout came with five new voices -- Arbor, Maple, Sol, Spruce, and Vale -- alongside improved accent handling and smoother conversations. The feature also launched on the web in November 2024, extending it beyond mobile. But Advanced Voice Mode still operated through discrete, alternating turns -- and it launched into the shadow of a PR debacle that OpenAI is still working to leave behind. The Scarlett Johansson controversy still shadows OpenAI's voice ambitions Advanced Voice Mode arrived in the wake of one of OpenAI's most damaging self-inflicted crises. During the GPT-4o launch in May 2024, the company showcased a voice called "Sky" that many listeners immediately noted sounded strikingly similar to Scarlett Johansson, who famously voiced an AI companion in the 2013 film Her. Johansson said she had declined OpenAI CEO Sam Altman's offer to voice the system, then was "shocked, angered and in disbelief" when the product launched with a voice her own friends couldn't distinguish from hers, as NBC News reported. Altman had tweeted just the word "her" the day the product launched. OpenAI pulled the voice and apologized, but the incident drew public scrutiny from SAG-AFTRA and members of Congress, and crystallized broader concerns about AI companies moving fast with creative IP. The Hollywood labor union said the issue underscored "why we're strongly championing federal legislation that would protect their voices and likenesses ... from unauthorized digital replication," as NBC News reported. Forbes contributor Paul Tassi wrote at the time that Altman, "by holding up Her on a pedestal of something to strive for, has missed the point of that film" -- in which the protagonist's relationship with his AI companion ultimately does him more harm than good. GPT-Live appears designed, in part, to move past those controversies. OpenAI says it has "remastered the nine distinct voices in ChatGPT for GPT-Live" and notes the system "is designed for conversation, not voice impersonation," with "safeguards to prevent it from imitating a real person's voice." What 150 million weekly voice users will actually notice today OpenAI disclosed that more than 150 million people talk to ChatGPT using voice and dictation features each week -- a notable slice of the platform's 900 million total weekly active users. The voice experience has grown into a substantial product in its own right, used for language practice, bedtime stories, commute-time chat, and hands-free everyday help. The new product features reflect that usage. GPT-Live introduces rich visual cards that surface during voice conversations -- weather forecasts, stock data, sports scores, and maps -- giving users something to glance at without breaking the flow of speech. Users can now choose between three reasoning levels for answers: Instant for quick responses, Medium for moderate thinking, and High for more complex work. And if you take a moment to think, "ChatGPT Voice now waits instead of jumping in and interrupting," OpenAI wrote. "If you ask it to stay quiet and listen, it will. And when there's background noise, like passing traffic or nearby conversations, ChatGPT is better at focusing on your voice instead of getting distracted." Early reactions from users with preview access were cautiously positive. "I had early access to sol. it is a phenomenal model," wrote one user on X, adding it is "much better at frontend, long context knowledge work, and its vibes are much better." Another observer cut to the heart of the matter: "The smarts are not new here, GPT-Live hands hard questions to GPT-5.5. What is new is the feel: full-duplex voice that listens while it talks." New voice-specific safety tests reveal where the risks still live The GPT-Live system card, published alongside the announcement, reveals a safety strategy built around the particular risks of real-time voice interaction -- a domain where the speed and intimacy of conversation create hazards that text-based chat does not. OpenAI expanded its safety evaluations to include audio-native tests, using both real user voice samples (from those who opted in) and synthetically generated prompts targeting edge cases across categories like self-harm, sexual content, illicit behavior, emotional reliance, mental health, and hate speech. On the synthetic evaluations -- which OpenAI described as deliberately adversarial -- GPT-Live-1 showed substantial improvements over Advanced Voice Mode. In illicit behavior, for instance, the safety score rose from 0.63 to 0.97. On self-harm, it climbed from 0.72 to 0.98. Hate speech achieved a perfect 1.00, up from 0.87. On the production-prompt evaluations -- which used real user audio and reflected more ambiguous, borderline scenarios -- the picture was more mixed. GPT-Live-1 matched or improved on Advanced Voice Mode in most categories but showed a slight regression on emotional reliance (from 0.88 to 0.82), though OpenAI noted the change was not statistically significant. The company built real-time safeguards that can intervene while the model is speaking -- steering toward safer responses, surfacing crisis resources, or ending the voice conversation entirely in higher-risk situations. It also designed additional protections for teen users and adapted self-harm support flows for voice, including crisis helpline integration. Perhaps most notably, OpenAI said it is "rolling out longer-term measurement and post-launch monitoring focused on emotional reliance" -- an acknowledgment that the very naturalness GPT-Live strives for creates its own category of risk. Google, ByteDance, and Nvidia are already in the full-duplex race While OpenAI was refining its safety guardrails, its rivals were shipping full-duplex systems of their own. Google's Gemini Live, which supports full-duplex conversation alongside camera and screen sharing -- capabilities GPT-Live notably lacks at launch -- is already available in the Gemini app. Google released Gemini 3.1 Flash Live in March as its highest-quality real-time audio model, targeting low-latency voice interactions for developers. ByteDance launched Seeduplex in April, claiming to be the first production-scale full-duplex speech AI deployed at scale, inside its Doubao app. Seeduplex reported roughly a 50 percent reduction in false-response and false-interruption rates compared to ByteDance's previous half-duplex system. And Nvidia's PersonaPlex, released in January, brought customizable voice and role control to full-duplex models, breaking what had been a constraint where natural-sounding models were locked into a single fixed voice. The competitive picture is clear: full-duplex voice interaction is quickly becoming table stakes for consumer AI products, not a differentiator. OpenAI's advantage lies in the scale of its existing user base, its integration with GPT-5.5's reasoning capabilities, and the breadth of the ChatGPT ecosystem. But the window in which any one company has a monopoly on natural-sounding voice AI has already closed. OpenAI also acknowledged several gaps. GPT-Live does not support voice with video or screen sharing at launch. Language support is limited, with the company noting that "for certain languages, the model may have a non-native accent or gaps in fluency." And API access is not available on day one, meaning enterprise developers cannot yet build on GPT-Live directly -- a constraint that will slow the model's penetration into commercial voice-agent workflows where competitors like Google, ElevenLabs, and Deepgram already have developer-facing products. The end of the chat box may be closer than anyone expected GPT-Live is essentially OpenAI's most significant bet yet on voice as the primary interface for AI -- not just a convenience feature bolted onto a text chatbot, but a purpose-built interaction layer that sits between the user and the company's most powerful models. "Over time, we believe this research will also unlock the ability to use voice for increasingly complex, longer-running, and more agentic work," OpenAI wrote. That ambition -- using natural voice as the front end for autonomous AI agents that can perform multi-step tasks -- is the logical endpoint of the full-duplex plus delegation architecture. Imagine telling your phone to book a flight, negotiate with your insurance company, or debug a production server, all through a conversation that feels as natural as talking to an assistant who also happens to have the intelligence of a frontier AI model. Two years ago, talking to ChatGPT meant dictating into a microphone and waiting nearly two seconds for a stilted reply. One year ago, it meant a smoother exchange that still felt like a polite, slightly awkward phone call with someone who insisted on waiting for you to finish every sentence. Today, it means something closer to a real conversation -- imperfect, still constrained in some languages and missing video, but unmistakably closer. OpenAI once got into trouble for wanting to recreate the movie Her. With GPT-Live, the company may finally be reckoning with the harder question the film actually posed: not whether AI can sound human enough to talk to, but what happens to us when it does.
[14]
OpenAI bets voice will become AI's primary interface with new models
Why it matters: The company sees this as a step toward a future where voice is the primary way people interact with AI. Zoom in: OpenAI says its smartest voice models yet, GPT-Live-1 and GPT-Live-1 mini, make conversations feel more human by allowing users to interrupt naturally and pause speech without the model cutting them off. * The voice models will now route queries to OpenAI's latest frontier text models, addressing what executives acknowledge has been a limitation for users. * They can also listen and speak at the same time, aiding in language translation capabilities. What they're saying: "This is the beginning. Over time, we think this will also unlock the ability to use voice as kind of the primary interface to computing," Atty Eleti, product lead for ChatGPT Voice, said. * It can also be used to "manage increasingly complex, long-running, agentic work." Follow the money: Voice is typically a more expensive and token-heavy way to interact with an AI model. * OpenAI said inference costs have come down enough to make the economics of voice work, so they're rolling this out to free users starting Wednesday. The intrigue: Given the company's push into hardware, the focus on voice could be an indication of future plans for a device that users interact with primarily via voice. * OpenAI said they did not have any hardware related news to share on a press briefing call with reporters about the voice models. Friction point: Audio is stored for 30 days for context and memory, though users can delete and export data any time.
[15]
OpenAI launches GPT-Live voice models for real-time conversations
OpenAI launched GPT-Live, a new family of voice models capable of listening and speaking at the same time, on Wednesday. The company said it is rolling out two versions -- GPT-Live-1 and GPT-Live-1 mini -- to users globally. GPT-Live-1 mini will replace the current Advanced Voice Mode in ChatGPT by default, while users on paid tiers will have access to the larger GPT-Live-1 model, the company said. Unlike traditional voice systems, the new models operate in full-duplex mode -- handling speech input and output at the same time -- which opens the door to natural interruptions and capabilities such as live translation, TechCrunch reported. Audio processing in the old system worked as a relay: spoken words were first transcribed, then fed to a language model, then converted back to speech before reaching the user. At a press briefing, OpenAI described shortcomings in the older system that the new models are meant to fix, pointing to cases where the assistant cut off users and fell short on question-answering. When a query calls for search, reasoning, or agentic capabilities, the system hands it off to a text model such as GPT-5.5 in the background without pausing the voice interaction, the company said. The new voice mode is also designed for longer conversations. ChatGPT Voice product lead Atty Eleti said during a press briefing that he had held 30- to 40-minute conversations with the feature during walks. Demonstrations also highlighted the model's ability to hold back and quietly take in what is being said until it is needed. Because the voice mode draws on newer GPT models, it can surface certain responses visually rather than through speech alone. The company said more than 150 million people use ChatGPT through voice features such as Voice and Dictation. OpenAI said it sees voice as a potential primary interface for complex computing work. "Over time, we think this will also unlock the ability to use voice as a kind of primary interface to computing, and to manage increasingly complex long-running agentic work," Eleti said. OpenAI said the models include safety features, such as guardrails that change responses for younger users and suggest support resources if sensitive topics like self-harm are mentioned. The company also said it is not presenting the new voice mode as an AI companion. Earlier this year in May, the company released a trio of audio models targeted at developers building voice-driven applications, with a focus on making those agents sound more natural in conversation, Reuters reported.
[16]
ChatGPT Live could make talking to AI feel straight out of the movies
AI voice assistants have been chasing the sci-fi dream for years, but they still have a hard time holding a conversation with humans. Most voice systems still need clear turns, clean pauses, and a few seconds before they respond. OpenAI is now rolling out GPT-Live, a new voice model for ChatGPT Voice that is designed to make those exchanges feel faster and less scripted. The main upgrade is what OpenAI calls a full-duplex architecture. In simpler terms, GPT-Live can listen and speak at the same time. It continuously processes what the user is saying while also generating its own response, allowing it to decide when to talk, when to pause, when to keep listening, and when to use a tool. What makes GPT-Live different from older voice systems? OpenAI says GPT-Live can handle the messier parts of a normal conversation more smoothly. Users can interrupt ChatGPT with another question, pause while thinking, ask it to slow down, or tell it to stay quiet and listen. It can also respond with small acknowledgments like "mhmm" or "got it," so the conversation does not feel as rigid. The model is also designed to work better in noisy environments. The company says ChatGPT Voice should now be better at focusing on the user's voice when there is background noise, such as traffic or nearby conversations. Recommended Videos GPT-Live can also hand off tougher requests in the background. If a question needs web search, deeper reasoning, or more complex work, it can delegate that task to GPT-5.5 and bring the result back when it is ready. The spoken conversation can continue while that happens. Who gets GPT-Live first? GPT-Live is rolling out globally in ChatGPT on iOS, Android, and the web. GPT-Live-1 will power ChatGPT Voice for Go, Plus, and Pro users. Free users will get GPT-Live-1 mini instead. OpenAI is also adding visual cards to voice chats. So, if you ask about weather, stocks, sports, or similar topics, ChatGPT can show a card on screen while continuing the voice conversation. Search, memory, images, and file uploads will also continue to work with ChatGPT Voice. Currently, the main limitation is around video. GPT-Live can handle voice conversations inside ChatGPT at launch, but it does not yet work with video or screen sharing. In other words, you cannot point your camera at something or share your screen during a GPT-Live voice chat yet. OpenAI says those features will arrive later.
[17]
I Tried ChatGPT's Improved Voice Mode, and It's More Natural Than Ever
* GPT-Live is the big new voice model upgrade for ChatGPT. * It's more natural and capable, and can listen while talking. * The upgrade is rolling out now across free and paid tiers. The latest upgrade being pushed out to ChatGPT, heading to all users now, is GPT‑Live. OpenAI is describing it as a "new generation" of voice models for interacting with the AI chatbot, and you might find that it leads you to spend more time chatting than typing. Voice mode for ChatGPT is nothing new, but previously it's been a relatively basic wrapper on top of the standard text input and output. It has been billed as a more natural way to engage with the AI, but GPT-Live promises to dial this fluidity up to an even higher level. For the first time, the voice mode will be able to think in the background while continuing the conversation. It'll also give you extra space to pause when you need it, and indicate it's still listening with phrases like "mhmm" or "yeah." You should find the upgrade on mobile and the web now (or very soon). Free users get access to GPT‑Live‑1 mini, while those on paid plans are able to access the even smarter GPT‑Live‑1 model. How GPT-Live works OpenAI's end goal is to make talking to ChatGPT feel like talking to a real person, and GPT-Live gets closer to that. Originally, interacting with the AI via voice required a specific model for speech-to-text, another for actually responding to the query, and another for text-to-speech. The previous voice mode in ChatGPT combined all of that into a single AI model, but it was still turn-based: You spoke, the chatbot answered, then you spoke again. With GPT-Live, ChatGPT can be talking and listening at the same time. You can interrupt it as and when needed, and responses should be faster and more nuanced. The new voice mode is supposedly smarter when it comes to recognizing the difference between you pausing mid-thought and actually finishing your query. The model now recalculates several times a second "whether to speak, continue listening, pause, interrupt, or invoke a tool." An added benefit of the upgrade is that even complex work and deep thinking can be passed back to ChatGPT's servers in the background, while the conversation is continuing. You can also tell ChatGPT to take a beat or slow down; visual responses have been improved as well, so you might, for example, see pop-up cards for locations, weather forecasts, and sports scores. You can also now ask GPT-Live to translate something into a foreign language as you speak. Thanks to the new capabilities, you'll hear a running translation in the other language as you talk, with no pauses or interruptions. Improvements have also been made in terms of ignoring background noise (like background traffic or conversations happening nearby). Testing out GPT-Live To get to voice mode in the mobile app, tap the soundwave-style icon to the right of the prompt box. The new mode looks a lot like the old one on the surface, but with this update, you should see Live at the top of the screen (for the time being, at least, you can tap this to switch back to the older models). Right away, the upgraded voice mode feels more realistic and natural. ChatGPT will talk in a varied and expressive way, throwing in useful markers like "let me check" whenever it's looking something up. It'll lso hesitate and draw words out at times. I chatted with GPT-Live for several minutes about upcoming movies, recent soccer matches, and tech news headlines, and got back answers that made sense and were respectfully brief (voice mode continues to be a refuge for those who don't want to see walls of text for every response). There were a couple of moments where the speech glitched and the conversation hung, but that was in about half an hour of chatting (presumably these bugs will get ironed out over time). Interruptions are handled well too, with the AI pausing to acknowledge what you've said and then continuing its train of thought. You can tweak the level of thinking ChatGPT puts into the new voice mode: Tap the sliders icon (top right), then tap Intelligence. There are three modes to pick from -- Instant, Medium, and High -- with varying levels of trade-off between the speed of the response and how detailed and accurate it is.
[18]
ChatGPT's 'smartest voice model ever' is rolling out to everyone today -- and GPT-Live-1 gives you more natural conversations without interruptions
OpenAI's new GPT-Live models are coming to everybody starting today * ChatGPT's new voice mode is rolling out today to everybody, even Free users * It allows for much more natural conversations and won't interrupt if you stop talking * You'll be able to do simultaneous translation for the first time ever in ChatGPT OpenAI has upgraded ChatGPT's voice mode for everybody with two new models that are rolling out globally, starting today. I listened to the new GPT-Live-1 model in a demo run by OpenAI, and it does sound much more natural than ChatGPT's previous voice model. The new model aims to address two particular problems with the existing ChatGPT voice mode. Firstly, the previous version just wasn't as smart as the text version of ChatGPT. Secondly, it tended to interrupt too much. You notice this especially if you go quiet while you're thinking of a reply -- ChatGPT will often fill the gap by talking. Sounding more intelligent To get around the intelligence problem, the new model actually delegates harder questions to ChatGPT-5.5, then comes back with an answer. It will say things like "let me just check that for you" to let you know it's doing this, which keeps the flow of conversation feeling natural and doesn't make it seem like you have to wait too long for an answer. It does the same thing with any answer it needs to look up on the web. So, for example, if you asked it when your team's next match was in the World Cup, it would say something like "OK, let me check that" while looking it up using GPT-5.5, then give you the answer. "Hey Chat" OpenAI also demonstrated how the new ChatGPT voice mode is quite happy to stop talking and listen if you tell it to, without interrupting. You can simply ask it not to reply until you speak to it directly again, and it will wait. Of course, this requires you to call it a name, which it doesn't officially have. In the demonstration I saw, the OpenAI employee called it "Chat", so he said "Hey Chat", just like you would say "Hey Siri". In practice that seems to work quite well. Simultaneous translation The final new feature of note is simultaneous translation. If you watch world leaders being briefed at places like the United Nations, you'll see that they have an earpiece through which they receive a simultaneous translation in their own language of whatever the speaker is saying. Now you can do this with ChatGPT. Say "I'd like you to simultaneously translate whatever I'm saying into [language]", then start talking, and ChatGPT will provide a live translation as you speak. Seeing this in action was actually quite impressive and I could imagine it being very handy in several real world situations. All major languages appear to be supported as well. The future for AI The new GPT-Live-1 models -- there are two, the normal one and a mini version -- will start rolling out for all users immediately, but it could take a few days to reach everybody. The smaller GPT-Live-1 mini model will be the default for Free users, while paid users get the full GPT-Live-1 model. So far, ChatGPT's voice mode has been a handy tool for when you need to use your hands for something and can't type, but it's never been good enough to become the standard way you interact with ChatGPT. Now it looks like OpenAI is trying to unlock the ability to use voice as the primary interface to AI, and it's quite possible that this is the future OpenAI is aiming for. Today I think we've all just taken a step closer to it. Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.
[19]
The viral influencer who broke ChatGPT's brain just proved OpenAI's latest GPT-Live model still can't beat him
OpenAI says GPT-Live is its most human AI voice model yet, but spelling simple vocab words still proves too big a challenge. ChatGPT just launched a new voice model, and above all else, its conversations are meant to feel natural. Seriously -- the word natural appears no less than 13 times in OpenAI's announcement of GPT-Live. Where previous voice models relied on turn-based conversations, GPT-Live aims to actively listen to users and continuously react, even interjecting with acknowledgments like "mmhmm" and "got it" to prove it's paying attention.
[20]
OpenAI launches GPT-Live voice model series ahead of broad GPT-5.6 release
OpenAI launches GPT-Live voice model series ahead of broad GPT-5.6 release OpenAI Group PBC today introduced GPT-Live, a family of artificial intelligence models optimized to process spoken instructions. The model series will power ChatGPT's voice mode. Additionally, OpenAI plans to make it available to developers via an application programming interface. GPT-Live includes two models on launch. GPT-Live-1, the more capable of the two, will be the default option in paid ChatGPT plans while a scaled-down algorithm called GPT-Live-1-mini will power the free tier. The models are rolling out to the iOS, Android and web versions of ChatGPT with support for multiple languages. ChatGPT's voice mode previously had to wait until users finished speaking before generating a response. As a result, it wasn't capable of performing real-time translation. Additionally, the feature could only perform tasks such as searching the web at the end of the user's monologue even if such tasks could theoretically be completed earlier. That slowed down prompt processing. According to OpenAI, GPT-Live features a so-called full-duplex architecture that addresses its predecessor's limitations. During conversations, the model series regularly checks whether it should interrupt the user, search the web or perform some other task. That removes the need for GPT-Live to wait until the user finishes describing a request. OpenAI developed a set of evaluations to study its new models' pleasantness. The company says that GPT-Live-1 scored 75.5, which put it well ahead of ChatGPT's previous voice processing model. When GPT-Live receives complex prompts, it routes them to GPT-5.5, OpenAI's most advanced commercially available model. The algorithm set records across several coding benchmarks when it rolled out in April. GPT-Live can use GPT-5.5 to browse the web, explain complex topics and perform a range of other tasks. OpenAI stated that GPT-5.5 will also be capable of using newer, more advanced reasoning models when they become available. The company reportedly plans to release three such algorithms later this week. Last month, OpenAI introduced a series of cutting-edge large language models called GPT-5.6. The most advanced LLM in the lineup, Sol, outperforms Claude Mythos 5 in some areas. It debuted alongside two scaled-down algorithms called Tera and Luna that trade off some output quality for lower pricing. Axios reported today that OpenAI will make GPT-5.6 broadly available on Thursday after receiving approval to do so from the U.S. government. Currently, GPT-5.6 is only accessible to a limited number of organizations. The White House's decision to authorize a broader release reportedly came after the Commerce Department's Center for AI Standards and Innovation carried out an evaluation of the model series. A few weeks before GPT-5.6's introduction, OpenAI formed a professional services business called the OpenAI Deployment Company. It focuses on helping organizations adopt the ChatGPT developer's LLMs. The venture has raised $4 billion in funding from OpenAI and 19 external backers. The OpenAI Deployment Company today announced that it's acquiring a fellow professional services provider called Northslope Inc. Its main specialty is building AI applications atop Palantir Technologies Inc.'s software. According to OpenAI, the company's revenue grew sevenfold last year.
[21]
GPT-Live finally gives ChatGPT the one thing Gemini already had
When not writing, Dave enjoys spending time with his family, running, playing the guitar, camping, and serving in his community. His favorite place is the Blue Ridge Mountains, and one day he hopes to retire there (hopefully his fear of heights will have retired by then, too!). * GPT-Live is an updated, advanced voice model for ChatGPT. * GPT-Live aims to be more conversational while handling deep queries in the background for a seamless feel. * A global rollout began on July 8 -- users on paid plans will get GPT-Live-1, and Free users will get GPT-Live-1 mini, a slightly scaled-back version. On July 8, OpenAI announced GPT-Live, an upgraded voice experience for ChatGPT. GPT-Live marks a major shift in ChatGPT's conversational capabilities, with a more natural feel and smarter responses. What is GPT-Live? Next-gen conversational AI GPT-Live is OpenAI's next-generation voice model. It aims to "make talking with AI feel much more like having a real conversation." GPT-Live is built on a "full-duplex architecture," which means it can listen and speak simultaneously. According to OpenAI, the model shows it's paying attention by using phrases like "mhmm" or "yeah" and engaging in back-and-forth conversation. It should also be able to tell when you're thinking and wait quietly for you to continue. The goal is "a voice experience that is refreshingly easy to talk to." OpenAI also claims GPT-Live is its "smartest voice model yet." It can delegate to the latest frontier models when necessary (GPT-5.5 at launch), and it can also continue the conversation while it thinks in the background, so conversations flow more naturally. Why GPT-Live is a big deal Playing catch-up This is a big step for ChatGPT's voice assistant. In the past, we've noted that other LLMs like Claude have felt more natural. And while the platform has had voice models for a while now, it's typically lagged behind competitors in this area, too -- especially Gemini. These new GPT-Live models should go a long way toward making ChatGPT more natural to use and help bring it up to par with Gemini Live and other competitors. Combined with features like Scheduled Tasks, ChatGPT could be a compelling alternative to Gemini. OpenAI also claims that GPT-Live is built with enhanced safeguards that can either steer responses to safer territory or end the conversation entirely in high-risk situations. There are "additional protections to support teen users," and the models are trained on age-appropriate behavior to tailor responses to the listener. GPT-Live availability Global rollout GPT-Live began rolling out globally on July 8. It's available across iOS, Android, and the web. Two models are being rolled out, depending on your subscription tier. Users on Go, Plus, and Pro will get access to GPT-Live-1, while Free users will get access to GPT-Live-1 mini (a similar but slightly scaled-down version). In both cases, GPT-Live will be the default model for ChatGPT Voice. The main limitations at this time are in language fluency and video. GPT-Live is optimized for "some of the most popular languages in ChatGPT," but OpenAI warns that certain languages may have a non-native accent or "gaps in fluency." The company says it's actively working on this. As for video, GPT-Live does not currently support voice with video or screen sharing. Again, OpenAI says it's working on this and will introduce these capabilities "soon." In the meantime, if you need these features, legacy versions of ChatGPT Voice that support them are still available. ChatGPT OS Android, iOS, Web Developer OpenAI Price model Free with optional subscription ChatGPT is OpenAI's flagship product, with powerful features. It was just updated with advanced, conversational voice capabilities. See at Google Play Store See at App Store Expand Collapse
[22]
OpenAI launches upgraded voice model GPT-Live that sounds more natural
OpenAI has introduced two new voice AI models for ChatGPT. These models enable simultaneous listening and speaking for natural interruptions. GPT-Live-1 mini is now the default, while GPT-Live-1 is for paid users. The new voice capabilities improve turn-taking and response speed significantly. Safeguards are included for age-appropriateness and self-harm topics. OpenAI has launched two new voice AI models, GPT-Live-1 and GPT-Live-1 mini, to make conversations with ChatGPT more natural. In a blog post, the company said the new models can listen and speak simultaneously, allowing users to interrupt naturally and use features such as live translation. They're also designed for longer conversations and faster responses.The company has made GPT-Live-1 mini the default instead of Advanced Voice Mode in
[23]
OpenAI Launches GPT-Live for Real-Time Conversations
Sam Altman's OpenAI has introduced a "GPT-Live" feature, a new-generation model that powers the ChatGPT voice, taking AI voice commands to the next level. The feature is exclusive to paid members and offers direct conversation with ChatGPT in a natural tone. * Make Telecom Talk My Trusted Source Also Read: JioHotstar Brings ChatGPT Powered Voice Discovery GPT-Live Launched: OpenAI's New AI Model for Live Conversations OpenAI brings this new update as an improved version of previous AI voice models, where tone and response depended on the user's input, making it more like a Siri-like assistant answering queries. Now, voice AI models have been upgraded to include new conversational-style models that interact with the user in a natural tone. With GPT-Live, ChatGPT can listen and speak simultaneously, enabling more natural back-and-forth, better timing, and live translation. Also Read: Swiggy Orders Can Now Take Place Through ChatGPT, Gemini and More ChatGPT's press release stated that the upgraded AI feature will have fewer interruptions during pauses and better handle background noise. This new ChatGPT voice feature also lets you choose intelligence levels: Instant, Medium, or High. The voice feature applies not only to live conversation but also to live translation. However, there is a catch! OpenAI has not enabled access to GPT-Live for everyone. According to the press release, OpenAI offers its new natural voice feature by default to Go, Plus, and Pro users. There is also a mini-version set as the default for Free users.
[24]
New ChatGPT Voice Full-Duplex Update Lets You Interrupt at Any Time
OpenAI's latest development, ChatGPT Voice, introduces full-duplex communication, allowing the system to listen and respond at the same time. This capability lets users refine their input mid-conversation without pauses, creating a more dynamic interaction. By incorporating contextual understanding, ChatGPT Voice can handle a variety of exchanges, from everyday conversations to addressing detailed or technical inquiries. Explore how ChatGPT Voice enables real-time translation to support multilingual communication with attention to cultural nuances. Learn about its language coaching feature, which provides tailored feedback on grammar and pronunciation for learners. Additionally, gain insight into how real-time search integration ensures responses remain accurate and up-to-date across diverse subject areas. Conversations Without Interruptions Full-Duplex Communication A defining feature of ChatGPT Voice is its full-duplex communication capability. Unlike traditional AI systems that alternate between listening and speaking, ChatGPT Voice can listen and respond simultaneously. This allows for more dynamic and natural interactions, where you can interrupt, pause, or adjust your input mid-conversation without disrupting the flow. For instance, if you're dictating a message and decide to rephrase halfway through, the system adapts instantly, maintaining a conversational rhythm. This innovation ensures smoother, uninterrupted exchanges, making the AI feel more like a genuine conversational partner. Advanced Reasoning: Thinking Beyond Responses ChatGPT Voice is designed to do more than provide simple replies, it reasons and analyzes. Equipped with advanced problem-solving capabilities, it can handle complex queries and deliver nuanced, context-aware answers. For tasks requiring deeper analysis, the system seamlessly integrates with GPT-5.5, a specialized model optimized for intricate problem-solving. Whether you're managing a multi-step project, troubleshooting technical issues, or brainstorming creative ideas, ChatGPT Voice offers intelligent, tailored solutions. This combination of conversational ease and robust reasoning improves the standard for AI-powered assistance, making it a valuable tool for both personal and professional use. Learn more about ChatGPT Voice with other articles and guides we have written below. Real-Time Translation: Breaking Language Barriers The real-time translation feature of ChatGPT Voice redefines cross-language communication by prioritizing meaning and cultural context over literal translations. This ensures that conversations remain natural and culturally appropriate. For example, if you're speaking in English but need to communicate with someone in French, ChatGPT Voice provides accurate and conversational translations that preserve the intended tone and meaning. This functionality is particularly beneficial for global collaboration, international travel and language learning, allowing seamless communication across diverse linguistic landscapes. Language Coaching: Your Personal Tutor For language learners, ChatGPT Voice acts as a personal language coach, offering real-time feedback on grammar, pronunciation and phrasing. If you make an error, the system gently corrects you and provides actionable suggestions for improvement. This interactive approach fosters confidence and fluency, making it an invaluable tool for mastering new languages. Whether you're preparing for a professional presentation or practicing casual conversations, ChatGPT Voice tailors its feedback to your specific needs, helping you achieve your language goals more effectively. Search and Reasoning Integration: Instant Access to Information ChatGPT Voice integrates real-time search capabilities, allowing it to retrieve and verify information instantly. If you ask a question that requires external data, the system conducts a web search and delivers concise, accurate answers. For example, whether you're discussing the latest news, exploring scientific concepts, or seeking restaurant recommendations, ChatGPT Voice combines its contextual understanding with up-to-date information to provide relevant insights. This feature ensures that you stay informed and engaged without needing to leave the conversation. User-Centric Features: Tailored to Your Needs ChatGPT Voice is designed with user-centric flexibility, allowing you to customize your interactions. You can choose between quick, concise exchanges or more detailed, in-depth discussions based on your preferences. This adaptability bridges the gap between voice and text-based AI models, making sure a consistent and personalized experience. By aligning with your communication style, ChatGPT Voice fosters a sense of trust and creates a more intuitive relationship between you and the AI, making it a versatile tool for various scenarios. Safety and Accessibility: Designed for Everyone Safety and accessibility are at the core of ChatGPT Voice's design. The system incorporates robust safeguards to prevent harmful or inappropriate interactions, making sure a secure environment for all users. Additionally, accessibility features make the technology inclusive for diverse needs. For instance, voice modulation options assist individuals with speech impairments, while simplified language settings cater to users who prefer straightforward communication. These measures enhance usability and inclusivity, making AI interactions feel more human, trustworthy and universally accessible. A Responsible Step Forward in Conversational AI The next generation of ChatGPT Voice represents a pivotal advancement in conversational AI. By integrating full-duplex communication, advanced reasoning, real-time translation, and user-centric features, it delivers a more fluid, intelligent and personalized interaction experience. Whether you're seeking assistance with daily tasks, solving complex problems, or breaking down language barriers, ChatGPT Voice is designed to meet your needs effectively. Its commitment to safety, accessibility and trust ensures that this innovation is not only advanced but also responsible, paving the way for a more inclusive and connected future in AI-driven communication. Media Credit: OpenAI Disclosure: Some of our articles include affiliate links. If you buy something through one of these links, Geeky Gadgets may earn an affiliate commission. Learn about our Disclosure Policy.
[25]
OpenAI Unveils GPT-Live As Competition Heats Up Around AI Voice Assistants - Microsoft (NASDAQ:MSFT)
Unlike traditional systems that wait for a user to finish speaking, GPT-Live listens and responds simultaneously. This allows it to handle instant interruptions, recognize pauses and provide brief conversational cues like "yeah" or "mhmm" for a human-like flow. The system also integrates directly with OpenAI's broader ecosystem. While handling standard voice interactions natively, GPT-Live automatically routes complex requests -- such as web searches or advanced reasoning -- to a separate frontier model before delivering the results seamlessly back through the audio interface. GPT-Live will initially rely on GPT-5.5 for heavier workloads and update as newer models become available. The company also noted that internal testing showed improvements over its previous Advanced Voice Mode across measures including conversational flow, interruption handling, and user preference during five- to 10-minute conversations. The update expands ChatGPT Voice beyond audio responses. OpenAI said the feature will begin showing visual cards during voice conversations for topics such as weather, stocks, and sports, while maintaining access to tools including search, memory, image generation, and file uploads. The new models include additional voice safety measures, including protections against unauthorized voice imitation and systems designed to handle sensitive conversations, using a limited set of approved voices rather than allowing users to create direct replicas of real people. The company acknowledged that performance may vary across languages, with some languages still experiencing differences in accent accuracy and fluency. "We've optimized GPT‑Live for some of the most popular languages in ChatGPT. For certain languages, the model may have a non-native accent or gaps in fluency. We're actively working to improve the experience across languages," the company said. AI Voice: The Next Major Interface The release marks another step in the race among AI developers, as companies compete to move beyond text-based chatbots and into more persistent, conversational assistants. Google has been expanding its Gemini-powered voice capabilities, Amazon is rebuilding Alexa around generative AI, and Apple is working to bring more advanced AI features into Siri through its Apple Intelligence platform. At the same time, startups including ElevenLabs are making a deeper push into AI-generated speech, voice agents, and conversational applications. In February, the company announced its most advanced text-to-speech model, Eleven v3. This content was partially produced with the help of AI tools and was reviewed and published by Benzinga editors. Market News and Data brought to you by Benzinga APIs To add Benzinga News as your preferred source on Google, click here.
[26]
OpenAI Introduces GPT-Live to Transform ChatGPT Voice with More Natural AI Conversations
The new model brings advanced voice interaction capabilities to ChatGPT, combining real-time communication, improved reasoning, live translation, and richer visual responses to create a more human-like conversational experience. OpenAI has unveiled GPT-Live, a next-generation voice model designed to make ChatGPT conversations more natural, responsive, and intelligent. The new model brings advanced voice interaction capabilities to ChatGPT, combining real-time communication, improved reasoning, live translation, and richer visual responses to create a more human-like conversational experience. GPT-Live represents a major evolution in OpenAI's voice technology, enabling users to interact with ChatGPT through smoother and more dynamic conversations. The model is designed to better understand context, respond faster, and handle complex discussions while maintaining a natural flow similar to human conversations. The improved architecture enables users to interrupt responses naturally, pause during conversations, and continue discussions without repeatedly restarting interactions. OpenAI said the technology is designed to deliver a more seamless experience by reducing delays and improving conversational responsiveness. This allows ChatGPT Voice to support more sophisticated use cases, including research assistance, professional discussions, learning support, and real-time decision-making. The model also includes enhanced listening capabilities, allowing it to better understand users in environments with background noise and reducing unnecessary interruptions caused by short pauses or changes in speech patterns. The updated ChatGPT Voice experience can provide visual response cards for topics such as weather updates, sports information, financial data, and other real-time information while continuing the voice conversation. The model also supports continuous real-time translation, enabling users to switch between languages more naturally. This capability is expected to improve communication across different languages by allowing conversations to continue without manual translation steps. As businesses and consumers increasingly adopt AI-powered tools, GPT-Live highlights the growing shift toward more interactive AI systems that combine voice, reasoning, translation, and real-time information access to deliver more intelligent digital experiences.
[27]
OpenAI Launches GPT Live Voice Models for Natural Human-AI Interaction
OpenAI's latest release, GPT Live, introduces a new chapter in AI voice technology with its focus on multitasking, real-time translation and natural conversational flow. Highlighting this development, AI Grid explores how GPT Live's full-duplex communication enables simultaneous listening and responding, creating a more fluid interaction compared to traditional voice models. For instance, users can issue commands or ask questions without pauses, making conversations feel uninterrupted and efficient. This feature, combined with its adaptability to noisy environments and overlapping speech, positions GPT Live as a practical solution for both casual and professional use. Dive into this feature to uncover how GPT Live's two-tiered approach, GPT Live 1 and GPT Live 1 Mini, caters to diverse needs, from advanced multitasking for professionals to accessible functionality for first-time users. Learn how its task delegation capabilities allow the AI to manage background operations like drafting emails or scheduling meetings without breaking conversational flow. Additionally, explore its integration of real-time translation and image analysis, which bridges language and visual barriers, offering new possibilities for collaboration and productivity. These insights showcase the versatility and potential of GPT Live in enhancing everyday interactions. Two Models, Tailored for Your Needs GPT Live is available in two versions, making sure accessibility and performance for a wide range of users: * GPT Live 1: The flagship model, equipped with innovative capabilities for users who require high-performance voice interactions. This version is ideal for professionals and advanced users seeking robust functionality. * GPT Live 1 Mini: A streamlined alternative designed for free-tier users, offering essential features without sacrificing quality. It provides an accessible entry point for casual users or those exploring AI voice technology for the first time. This tiered approach ensures that GPT Live meets diverse requirements, making it suitable for both personal and professional applications. Full-Duplex Communication: Conversations Without Interruptions One of the most notable features of GPT Live is its full-duplex communication capability. Unlike traditional voice models that alternate between listening and responding, GPT Live can process input and output simultaneously. This creates a smooth, uninterrupted conversational flow, allowing you to engage in natural, real-time interactions. Whether you're issuing commands, asking questions, or holding detailed discussions, this feature ensures efficiency and ease of use. Learn more about ChatGPT by reading our previous articles, guides and features : Task Delegation: Multitasking Made Simple GPT Live 1 introduces task delegation, a feature that enables the AI to perform complex operations in the background while maintaining an active conversation. For example, you can discuss a project while the system drafts an email, schedules a meeting, or conducts research, all without interrupting the dialogue. This functionality enhances productivity by streamlining multitasking, making it easier to manage multiple responsibilities simultaneously. Enhanced Accuracy and Adaptability The GPT Live models deliver significant improvements in accuracy, particularly in challenging scenarios such as overlapping speech or noisy environments. Benchmarks like GPQA and agentic search demonstrate these advancements, making sure reliable performance across various contexts. Whether you're engaging in casual conversations or handling high-stakes professional tasks, GPT Live adapts to your needs, providing consistent and dependable results. Real-Time Translation: Bridging Language Gaps GPT Live excels in real-time translation, allowing seamless communication across different languages. The AI can translate live conversations while preserving a natural tone, making it an invaluable tool for international collaboration, travel and multilingual environments. This feature eliminates language barriers, allowing you to connect with others effortlessly, regardless of linguistic differences. Temporal Awareness: Time-Sensitive Assistance With built-in temporal awareness, GPT Live can manage time-sensitive tasks effectively. Whether you need to set timers, schedule reminders, or track deadlines, the AI remains contextually aware of time within conversations. This capability adds practicality to interactions, making GPT Live a reliable assistant for organizing your day and staying on top of your commitments. Image Analysis Integration: Voice Meets Visual Context GPT Live goes beyond voice interactions by incorporating image analysis capabilities. Using voice commands, you can discuss and analyze images, allowing dynamic interactions that combine visual and auditory inputs. This feature is particularly beneficial for creative professionals, educators and individuals working with visual data, offering a new dimension to AI-assisted tasks. By bridging the gap between voice and visual context, GPT Live enhances its versatility and utility. Boosting Productivity Through Integration GPT Live integrates seamlessly with mini-browsers and apps, transforming it into a powerful productivity tool. It can perform tasks traditionally limited to text input, such as conducting research, managing calendars, or providing contextual assistance during live conversations. This integration ensures that GPT Live adapts to your workflow, whether you're managing a complex project, organizing daily activities, or seeking real-time support. Its ability to connect with external tools enhances its practicality and effectiveness. Versatility Across Applications The flexibility of GPT Live makes it suitable for a wide array of use cases: * Real-time translation: Communicate effortlessly across languages in live conversations. * Task management: Delegate, track and complete tasks with ease. * Contextual assistance: Receive relevant support tailored to your ongoing discussions. * Creative collaboration: Analyze images and brainstorm ideas using voice commands. Whether you're collaborating with colleagues, learning a new language, or streamlining your daily schedule, GPT Live adapts to your needs, enhancing both personal and professional experiences. Redefining AI Voice Technology OpenAI's GPT Live represents a significant advancement in AI voice technology. By addressing key challenges such as multitasking, contextual understanding and real-time translation, it delivers a user experience that feels intuitive and efficient. With its innovative features, adaptability and focus on practical applications, GPT Live sets a new standard for AI voice interactions. It paves the way for a future where communicating with AI is as seamless and natural as conversing with another person. Media Credit: TheAIGRID Disclosure: Some of our articles include affiliate links. If you buy something through one of these links, Geeky Gadgets may earn an affiliate commission. Learn about our Disclosure Policy.
[28]
OpenAI launches GPT-Live voice models for ChatGPT By Investing.com
Investing.com -- OpenAI released GPT-Live on Wednesday, a new generation of voice models designed to make conversations with artificial intelligence more natural. The models now power ChatGPT Voice. GPT-Live uses a full-duplex architecture that allows it to listen and speak simultaneously. During conversations, the model can acknowledge input with phrases like "mhmm" or "yeah," engage in rapid exchanges, or remain silent when needed. The company said GPT-Live can delegate tasks requiring web search, deeper reasoning, or complex work to frontier models running in the background. At launch, GPT-Live will use GPT-5.5 for these tasks. OpenAI plans to update the underlying model as new frontier models become available. OpenAI is rolling out two versions: GPT-Live-1 and GPT-Live-1 mini. Both are being distributed to ChatGPT users globally starting Wednesday. The company said it plans to make the models available through its API soon. Previous voice systems relied on cascaded models that processed speech in sequence, or turn-based models that required users to finish speaking before responding. GPT-Live processes input continuously while generating output, allowing it to make interaction decisions multiple times per second. In head-to-head evaluations, GPT-Live-1 and GPT-Live-1 mini were preferred over Advanced Voice Mode in conversations lasting five to ten minutes. The evaluations measured overall preference, turn-taking, interruptions, conversational flow, and how natural interactions felt. More than 150 million people use ChatGPT Voice and Dictation features weekly, according to OpenAI. The new ChatGPT Voice experience includes natural conversations, smarter answers with access to frontier models, improved listening capabilities, and visual responses. Users can choose reasoning levels: Instant for fast responses, or Medium and High for more complex thinking. ChatGPT Voice can now display visual cards for topics including weather, stocks, and sports while users are talking. The feature continues to support search, memory, images, and file uploads. GPT-Live-1 will become the default model for Go, Plus, and Pro users, while GPT-Live-1 mini will serve Free users. The models are available on iOS, Android, and ChatGPT.com. At launch, GPT-Live will not support voice with video or screen sharing in ChatGPT. OpenAI said it is working to add these capabilities. Users can still access previous versions of ChatGPT Voice where these features are available. This article was generated with the support of AI and reviewed by an editor. For more information see our T&C.
[29]
ChatGPT's voice chat gets an update, and it's simply impressive
GPT-Live lets ChatGPT listen, speak, and translate simultaneously OpenAI has unveiled GPT‑Live, its new voice model, which completely changes the way we talk to ChatGPT. Now we can have much more natural and fluid conversations with the AI and get the feeling that we're really talking to a person. The new system is so surprising that, from the very first use, it's clear we've moved to a new level of interaction. Natural conversations without pauses The key to GPT‑Live lies in its full‑duplex communication architecture, capable of listening and speaking at the same time. This means we can talk to it while it responds, interrupt it, or let it give us a few seconds to think, without cuts or artificial silences. It even uses expressions like "mhmm" or "I see" to keep the rhythm of the conversation and show us that it's still there. During the presentation, GPT‑Live demonstrated, for example, its simultaneous translation capabilities, in which ChatGPT knew to wait until the original sentence had enough meaning and context to do the translation, instead of going phrase by phrase or, worse, word by word. It was also striking how it's able to interrupt while a person is practicing speaking English to make relevant corrections. The naturalness of the interaction, even its laughter at certain points, is simply spectacular. Available now for all users The new version is already available in ChatGPT on iOS, Android, and the web, with GPT‑Live‑1 for those of us on the Go, Plus, and Pro plans, and GPT‑Live‑1 Mini for those using the free version. In the next few days it will also be available in the API for developers. With GPT‑Live, talking to ChatGPT changes completely. A much simpler and more fluid interaction that invites you to use it at any time of day.
Share
Copy Link
OpenAI introduced GPT-Live-1 and GPT-Live-1 mini, new conversational voice models that can speak and listen at the same time, replacing Advanced Voice Mode in ChatGPT. The full-duplex models handle interruptions naturally, delegate complex tasks to GPT-5.5, and aim to make voice the primary interface for computing. Over 150 million people already use ChatGPT's voice features.
OpenAI has released new voice models for ChatGPT called GPT-Live-1 and GPT-Live-1 mini, marking a shift in how users interact with conversational AI models
1
. The company is replacing its current Advanced Voice Mode with GPT-Live-1 mini by default, while paid users on Go, Plus, and Pro tiers gain access to the larger GPT-Live-1 model4
. These models represent a fundamental architectural change from previous systems that combined separate speech-to-text, language processing, and text-to-speech components into a unified experience.
Source: SiliconANGLE
The breakthrough feature of these new voice models lies in their full-duplex architecture, which allows ChatGPT to speak and listen at the same time
2
. This continuous interaction framework means the AI can receive information and produce outputs simultaneously, enabling users to interrupt naturally without derailing the conversation3
. During conversations, GPT-Live shows it's paying attention with phrases like "mhmm," "yeah," or "got it," creating more natural live conversations that mirror human-like conversations5
. The model can also stay silent for extended periods, understanding long pauses in speech as moments when users need time to think rather than prompts to start responding.
Source: VentureBeat
When GPT-Live encounters queries requiring research, reasoning, or agentic capabilities, it delegates tasks to OpenAI's frontier GPT-5.5 model while maintaining conversational fluidity
1
. Kundan Kumar, research lead for the GPT-Live model, explained that "when GPT-Live has to think hard for a question, it can delegate its reasoning and complex task to GPT-5.5, which can do things in parallel, and this GPT-Live can still remain in conversation with the user"2
. Users can select from three intelligence levels—Instant for quick responses, Medium for deeper analysis, and High for complex problem solving—depending on their needs3
.OpenAI envisions voice as the future primary interface to computing for increasingly complex long-running agentic work
1
. ChatGPT Voice's product lead, Atty Eleti, reported having 30- to 40-minute-long conversations with the voice feature during walks, demonstrating the system's capability for extended interactions1
. More than 150 million people already talk to ChatGPT using features like Voice and Dictation, indicating substantial existing demand1
. The models can also present information in visual formats when appropriate, with rich visual cards appearing for topics like weather, stocks, and sports4
.
Source: TechRadar
Related Stories
OpenAI emphasized expanded safety testing to address concerns about anthropomorphizing AI, particularly for users struggling with mental health issues
2
. The new models include safeguards to provide age-appropriate responses to teens and resources if conversations turn to topics like self-harm, psychosis, violence, and sexual content1
. OpenAI automatically opts users out of AI training with voice mode, storing audio clips for 30 days to maintain context from prior conversations, with deletion options available2
. Despite making conversations feel more natural, the company stressed it's not aiming to create an AI companion.The GPT-Live models are rolling out now to all ChatGPT users worldwide on iOS, Android, and web platforms, though OpenAI notes it may take a day or two for everyone to gain access
2
. Free users receive GPT-Live-1 mini, while paid users access the more capable GPT-Live-1 version4
. Legacy models like Standard and Advanced Voice Mode remain accessible from the app for users who prefer them4
. The models have been optimized for most spoken languages, though during demonstrations, the live translation feature in Hindi exhibited a heavy American accent and unnatural, bookish-sounding speech1
.Summarized by
Navi
[5]
1
Technology

2
Policy and Regulation

3
Technology
