3 Sources
[1]
Microsoft targets ultra-realistic voice agents with its first streaming transcription model
Microsoft targets ultra-realistic voice agents with its first streaming transcription model Microsoft Corp. today expanded its MAI artificial intelligence model family with its first streaming transcription model, debuting alongside two others focused on text-to-speech. They're designed for
[2]
Microsoft Targets Contact Center Market With Real-Time AI Voice Models | PYMNTS.com
"Together, they give developers more choice in how to build natural voice experiences," Naomi Moneypenny, senior director of product development, Microsoft Foundry Models, said in the post. One model, MAI-Transcribe-2-Streaming, turns live speech into text as it arrives. This model transcribes
[3]
Microsoft launches new voice agents: MAI-Transcribe-2-Streaming and MAI-Voice-2.1 explained
Have you ever noticed how uncomfortable that silence is when communicating with a voice bot? After you have completed your thought and there is nothing else you have to say, the bot listens to you until you stop, converts everything into text form, thinks, and after all this, speaks. Microsoft AI
Share
Copy Link
Microsoft AI released three new voice models—MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash—designed to enable ultra-realistic voice agents that respond in near real-time. The streaming transcription model delivers results within 320 milliseconds across 60 languages, while the text-to-speech models support 23 languages with natural cross-language voice consistency.
Microsoft AI expanded its MAI model family with three new releases aimed at building ultra-realistic voice agents capable of natural, human-like conversations. The company introduced MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash—models designed to eliminate the awkward pauses that plague current voice bots by processing speech and generating responses in near real-time
1
2
. These AI voice models represent a strategic shift as Microsoft AI reduces dependency on external providers like OpenAI and Anthropic, instead developing proprietary technology to power conversational AI agents across its ecosystem1
.
Source: PYMNTS
MAI-Transcribe-2-Streaming marks Microsoft's first streaming transcription model, accepting human speech via WebSocket and transforming it into continuously updated transcripts as people speak
1
. The live speech-to-text model delivers its first transcript hypotheses within 320 milliseconds on average, enabling applications to begin processing user requests before they finish speaking1
. This low latency capability allows customer service agents to identify caller requests mid-sentence and voice assistants to start reasoning or preparing tool calls sooner2
.
Source: SiliconANGLE
The model supports more than 60 languages with automatic language detection, making it suitable for global deployment
1
3
. Microsoft claims the model ranks number one on Artificial Analysis for accuracy in both final and partial transcript forms, producing words twice as fast as competitors3
. Priced at $0.54 per audio hour through Microsoft Foundry and Vercel AI Gateway, the streaming variant costs significantly more than the non-streaming MAI-Transcribe-2 model at $0.10 per audio hour due to continuous processing requirements1
3
.MAI-Voice-2.1 delivers high-fidelity outputs as Microsoft's most expressive text-to-speech model, generating natural speech across 23 languages and 26 locales
2
3
. The model maintains cross-language voice consistency, allowing the same voice to speak multiple languages with natural accents rather than applying one accent across all languages3
. Priced at $22 per million characters, MAI-Voice-2.1 is designed for applications where voice quality is central to the user experience1
2
.MAI-Voice-2.1-Flash optimizes for speed and volume, offering the same language support and cross-language voice identities as its counterpart but prioritizing responsive, high-volume voice applications
1
2
. With end-to-end latency of 150 milliseconds and 55% faster inference speed, Flash costs $15 per million characters—60% lower than competitive options3
. Both models support voice cloning from 5 seconds of audio reference with built-in consent guardrails3
.Related Stories
The releases demonstrate Microsoft's accelerating transition from model providers like OpenAI and Anthropic, despite major investments in both companies
1
. Microsoft AI Chief Executive Mustafa Suleyman has expressed concern about costs associated with frontier models, instructing researchers to focus on the MAI model family with the goal of powering Copilot agents in Excel and Outlook1
. "We pay a lot of money to Anthropic, so our goal is to reduce and ultimately eliminate that cost," Suleyman told Bloomberg1
.Microsoft is targeting the contact center market with these real-time AI voice models, enabling customer service agents to handle interactions more naturally
2
. The company announced in July a $2.5 billion investment in the Microsoft Frontier Company, a new unit aimed at helping customers implement AI by embedding 6,000 industry and engineering experts2
. This move addresses the gap between ideation and integration as enterprise customers seek practical AI implementations2
.Voice agents require three core capabilities: recognizing and understanding speech, deciding appropriate actions, and generating audible responses
1
. MAI-Transcribe-2-Streaming handles speech recognition, MAI-Voice models generate responses, and Microsoft's reasoning model Mai-Thinking-1 processes transcripts to determine agent actions1
. This modular approach gives developers control over quality, latency, and costs when building conversational AI agents1
.All three models are available through Microsoft Foundry, MAI Playground, Vercel, Azure Voice Live, and OpenRouter, with LiveKit integration coming soon
3
. Microsoft created a demo project called Chatter on the MAI Playground showcasing multilingual agents that respond in the language of incoming messages, tutors, and role-play scenarios requiring multiple speakers3
.Summarized by
Navi
[1]
02 Apr 2026•Technology

29 Aug 2025•Technology

08 May 2026•Technology
