4 Sources
[1]
Meta's new AI transcription model can distinguish between multiple speakers and languages in real-time - Engadget
Meta has introduced its first real-time audio model, Muse Voice Transcribe. The model can handle dictation and transcription for more than 20 speakers and can seamlessly handle multiple languages at once, Meta says. Meta CEO Mark Zuckerberg, who recently returned to X after three years of not
[2]
Meta Muse Voice Transcribe Supports 5 Indian Languages, 70+ Globally
* Muse Voice Transcribe supports more than 70 languages * It can distinguish more than 20 speakers * The model can transcribe code-switched conversations Meta has launched Muse Voice Transcribe with native support for five major Indian languages, including Hindi, Tamil, Telugu, Kannada and
[3]
Meta Launches Muse Voice Transcribe with Multilingual AI Support
Meta introduced Muse Voice Transcribe, a new AI model designed to listen, transcribe, and distinguish between speakers in real time. The model supports streaming transcription, can identify more than 20 speakers in a conversation, and can detect when a person finishes speaking while audio is being
[4]
Meta launches Muse Voice Transcribe with real-time speaker detection and multilingual support
Meta is rolling out the model across Meta AI, Muse Code, its Model API and Meta AI for Mac. Meta has introduced Muse Voice Transcribe, a new artificial intelligence model designed to listen, transcribe and distinguish between speakers in real time. It can handle streaming transcription, identify
Share
Copy Link
Meta Superintelligence Labs unveiled Muse Voice Transcribe, its first real-time audio perception model that combines live speech transcription, speaker separation, and endpointing in a single system. The AI model distinguishes between more than 20 speakers, handles code-switching across 70+ languages, and processes audio in 80-millisecond chunks using adaptive delay technology.
Meta Superintelligence Labs has launched Meta Muse Voice Transcribe, marking the company's entry into real-time audio transcription with an AI model that processes speech as it happens rather than after recording completion
1
. The real-time transcription system combines live speech transcription, speaker separation, and endpointing within a single architecture, eliminating the need for separate post-processing steps2
. Mark Zuckerberg demonstrated the model's capabilities on X, showcasing its ability to automatically distinguish between multiple speakers while seamlessly switching between languages mid-conversation1
.
Source: Gadgets 360
The real-time audio transcription model can identify and track more than 20 speakers in recordings longer than an hour, making it suitable for complex, multi-participant conversations
2
. Meta trained the system on over 70 languages, with 25 undergoing extensive validation for the initial release, including Chinese, French, Hindi, Japanese, Spanish, and Vietnamese3
. The model provides native support for five major Indian languages—Hindi, Tamil, Telugu, Kannada, and Malayalam—expanding accessibility across diverse linguistic markets2
. Real-time speaker detection happens as audio streams in, allowing the system to identify when someone starts or stops talking without delay4
.Muse Voice Transcribe processes incoming audio in 80-millisecond chunks, using adaptive delay to determine when sufficient information exists to transcribe each word accurately
1
. The AI model commits faster on simpler words while waiting slightly longer on complex phrases, balancing speed with accuracy through this intelligent processing approach4
. Meta employed reinforcement learning during training to optimize the trade-off between transcription accuracy and lower latency, helping the system make split-second decisions about when to finalize each transcribed token3
. This adaptive delay system allows the model to handle messy, real-world audio conditions while maintaining high accuracy across hour-long sessions1
.The model excels at handling code-switching, where speakers mix multiple languages within a single sentence—a common occurrence in multilingual conversations
1
. Users don't need to manually change language settings when speakers switch between languages, as the system automatically detects and adapts to linguistic shifts in real time2
. Meta demonstrated this capability by showing the model switching between English and Mandarin within the same sentence, maintaining accuracy throughout the transition3
. The system supports language, keyword, and context biasing, helping it identify words using information from both the audio stream and the broader conversational context2
.Related Stories
Meta has made Muse Voice Transcribe available through the Meta Model API at $3 per 1,000 audio minutes, which translates to approximately $0.18 per hour of transcription
2
. The model is immediately rolling out across several Meta products, including voice dictation in Meta AI, Muse Code, and Meta AI for Mac4
. Because the Meta AI for Mac app can power voice-enabled features across other applications, Muse Voice Transcribe will now handle dictation features on third-party services1
. Developers can access the model to integrate speech transcription capabilities into their own applications, with a demo version available on Meta's research blog1
.
Source: Engadget
Meta's release follows Google's launch of Gemini 3.5 Transcribe by less than a week, intensifying competition in the real-time audio transcription space
1
. While Google integrates its audio model into Android and Chrome, Meta's integration strategy remains focused on its own ecosystem and developer access through APIs1
. Meta positions Muse Voice Transcribe as a building block for personalized AI assistants capable of understanding accents, interruptions, overlapping speech, and multilingual conversations—capabilities increasingly demanded in modern voice interfaces3
. The model represents the latest release from Meta Superintelligence Labs, which has recently introduced its first dedicated coding agent, an open-weight model, and the Meta AI Mac app1
. As real-time speech systems evolve beyond basic transcription, speaker identification and faster response times are becoming critical differentiators in AI-powered voice tools3
.Summarized by
Navi
[3]
1
Technology

2
Technology

3
Policy and Regulation
