7 Sources
[1]
OpenAI upgrades its transcription and voice-generating AI models | TechCrunch
OpenAI is bringing new transcription and voice-generating AI models to its API that the company claims improve upon its previous releases. For OpenAI, the models fit into its broader "agentic" vision: building automated systems that can independently accomplish tasks on behalf of users. The
[2]
OpenAI's new voice AI model gpt-4o-transcribe lets you add speech to your existing text apps in seconds
Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More OpenAI's voice AI models have gotten it into trouble before with actor Scarlett Johansson, but that isn't stopping the company from continuing to advance its offerings in
[3]
OpenAI's new voice AI can apologize like it actually means it
According to TechCrunch, OpenAI is launching upgraded transcription and voice-generating AI models in its API, which the company claims enhance prior versions. This release aligns with OpenAI's broader aim of creating automated systems that can autonomously perform tasks for users. The new
[4]
OpenAI Just Released Its Latest Voice AI Tech, and It's Highly Customizable
OpenAI is releasing new text-to-speech and speech-to-text AI models, which it says will enable developers to build AI agents with highly customizable voices. The release of the audio models comes just over a week after OpenAI debuted a new application programming interface (API) and software
[5]
OpenAI's New Audio Models in API Can Be Used to Build Speaking AI Agents
OpenAI's new generation of audio models outperforms its existing models OpenAI, on Thursday, introduced new audio models in application programming interface (API) that offer improved performance in accuracy and reliability. The San Francisco-based AI firm released three new artificial
[6]
OpenAI Launches New Speech-to-Text AI Audio Models API for Developers
OpenAI has today introduced a suite of advanced audio models and tools through its API, designed to empower developers in creating sophisticated, voice-driven applications. These updates include innovative speech-to-text and text-to-speech models, seamless integration via the Agents SDK, and tools
[7]
OpenAI AI Audio : TTS Speech-to-Text Audio Integrated Agents
OpenAI has introduced a series of AI audio models, fundamentally redefining how voice-based AI can be integrated into modern applications wit&h ChatGPT. These advancements include state-of-the-art speech models, enhanced APIs, and comprehensive tools for developing voice agents. By focusing on
Share
Copy Link
OpenAI introduces new AI models for speech-to-text and text-to-speech, offering improved accuracy, customization, and potential for building AI agents with voice capabilities.

OpenAI has unveiled a new suite of AI models designed to revolutionize speech-to-text and text-to-speech capabilities. These models, integrated into OpenAI's API, promise enhanced accuracy, customization, and the potential to build more sophisticated AI agents with voice interactions
1
2
.The company has introduced two new speech-to-text models: gpt-4o-transcribe and gpt-4o-mini-transcribe. These models are set to replace OpenAI's previous Whisper model, offering significant improvements in transcription accuracy
1
.Key features of the new transcription models include:
1
3
Jeff Harris, a member of OpenAI's product staff, emphasized the importance of accuracy: "Making sure the models are accurate is completely essential to getting a reliable voice experience"
1
.The new text-to-speech model, gpt-4o-mini-tts, introduces enhanced "steerability" and customization options
1
2
. Developers can now:1
4
These audio models align with OpenAI's broader vision of creating "agentic" AI systems capable of independently accomplishing tasks
1
. The company recently released an Agents SDK, allowing developers to incorporate voice interactions into existing text-based applications with minimal code changes2
5
.The new models are available through OpenAI's API with the following pricing structure:
5
Related Stories
These advancements come at a time of increasing competition in the AI transcription and speech space. Companies like ElevenLabs and Hume AI are offering their own specialized models with unique features such as diarization and word-level customization
2
.Unlike its predecessor Whisper, OpenAI has chosen not to make these new transcription models openly available. The company cites the models' increased size and complexity as reasons for this decision, stating that they are not suitable for local execution on personal devices
1
3
.As AI continues to evolve, OpenAI's latest audio models represent a significant step forward in creating more natural and versatile voice interactions, potentially transforming various industries from customer service to creative storytelling.
Summarized by
Navi
[2]
1
Science and Research

2
Policy and Regulation

3
Technology