4 Sources
[1]
Meta's new AI transcription model can distinguish between multiple speakers and languages in real-time - Engadget
Meta has introduced its first real-time audio model, Muse Voice Transcribe. The model can handle dictation and transcription for more than 20 speakers and can seamlessly handle multiple languages at once, Meta says. Meta CEO Mark Zuckerberg, who recently returned to X after three years of not posting on the platform, shared an example of the model's ability to handle multiple speakers and languages at once. In the video, the transcription is able to automatically distinguish between multiple speakers and switch between languages. It's even able to pick up on "code-switching" and transcribe sentences that use words from multiple languages. "The model decides when to listen. It waits a little longer on hard words and commits faster on easy ones, using adaptive delay to predict each token and increase accuracy," he explained. "It holds up on messy, real audio too -- trained across 70+ languages (with 25 validated at launch), handles mid-sentence code-switching, and manages hour-long sessions with 20+ speakers. Meta's release comes less than a week after Google Gemini 3.5 Transcribe, its own audio model that boasts similar capabilities. But while Google is baking its audio model into Android and (eventually) Chrome, it's not clear if Meta has plans to integrate Muse Voice Transcribe into its flagship services. For now though, people can experience the new mode's capabilities in Meta's recently released Meta AI Mac app. Because the Mac app is able to power voice-enabled features on other apps, Muse Voice Transcribe will now power dictation features on other services. The model is also available to developers within Muse Code and Meta's Model API. It's priced at $3 for 1,000 audio minutes. Additionally, theres a demo version of the Muse Transcribe on Meta's research blog. Muse Voice Transcribe is the latest release from Meta Superintelligence Lab (MSI), which has been churning out new AI models and tools. In the last few weeks, the company has also introduced its first dedicated coding agent, an open-weight model and the aforementioned Meta AI Mac app.
[2]
Meta Muse Voice Transcribe Supports 5 Indian Languages, 70+ Globally
* Muse Voice Transcribe supports more than 70 languages * It can distinguish more than 20 speakers * The model can transcribe code-switched conversations Meta has launched Muse Voice Transcribe with native support for five major Indian languages, including Hindi, Tamil, Telugu, Kannada and Malayalam, as part of support for more than 70 languages globally. The model is the first real-time audio perception model developed by Meta Superintelligence Labs. It combines live speech transcription, speaker separation and endpointing in a single system, while also handling code-switching and conversations with more than 20 speakers. Meta Muse Voice Transcribe Brings Real-Time Transcription Muse Voice Transcribe can turn speech into text as someone is speaking, rather than processing the entire recording afterwards, the company explained in a blog post. It can also identify individual speakers and detect when someone starts or stops talking. Transcription, speaker separation and endpointing happen as the audio comes in, without a separate post-processing step. It can distinguish more than 20 speakers in a recording and work with audio longer than an hour, according to Meta. This allows it to handle long conversations without limiting the transcription to shorter clips. The model processes audio in 80-millisecond chunks and determines when it has heard enough information to transcribe each word. It can wait longer for words that are harder to recognise and commit simpler words sooner. Meta says reinforcement learning helps the model decide these delays while keeping transcription errors low. Meta trained the model on more than 70 languages and extensively validated 25 of them for the initial release. Its language support includes five major Indian languages, including Hindi, Tamil, Telugu, Kannada and Malayalam, the company added in a press release. It can also transcribe conversations in which speakers switch between languages, including when the switch happens within a sentence. Users do not have to change the language manually each time someone switches languages. The model supports language, keyword and context biasing, which helps it identify words using information from the audio and the wider conversation. Meta Muse Voice Transcribe Availability and Pricing Meta has made Muse Voice Transcribe available through the Meta Model API, so developers can use the model for speech transcription in their own applications. Meta AI for Mac and Muse Code already use it for dictation. The API is priced at $3 (roughly Rs. 300) per 1,000 audio minutes, which works out to around $0.18 (roughly Rs. 17) per hour.
[3]
Meta Launches Muse Voice Transcribe with Multilingual AI Support
Meta introduced Muse Voice Transcribe, a new AI model designed to listen, transcribe, and distinguish between speakers in real time. The model supports streaming transcription, can identify more than 20 speakers in a conversation, and can detect when a person finishes speaking while audio is being captured. Meta says Muse Voice Transcribe can also switch between languages during a conversation. The model can be fine-tuned to recognize specific names, places, and other terms, adding another layer of flexibility for transcription tasks. Meta stated that Muse Voice Transcribe processes incoming audio in 80-millisecond slices. It uses an adaptive delay system to determine when enough information is available to transcribe a word. Easier segments are processed almost instantly, while more complex phrases receive slightly more processing time. The company trained the system using reinforcement learning, balancing transcription accuracy with lower latency. Meta claims this approach helps improve the trade-off between speed and accuracy. The same architecture also handles and detects conversational endpoints. This allows transcription and speaker tracking to operate within a single system. Also Read: According to Meta, Muse Voice Transcribe was trained on over 70 languages, with 25 languages undergoing extensive validation. These include Chinese, French, Hindi, Japanese, Spanish, and Vietnamese. Meta also demonstrated the model switching between English and Mandarin within the same sentence. The company showcased its ability to process long recordings involving multiple speakers without requiring manual cleanup. Meta sees Muse Voice Transcribe as a building block for more personalized AI assistants that can understand accents, interruptions, overlapping speech, and multilingual conversations. The model is rolling out immediately across several Meta products, including voice dictation in Meta AI and the company's Muse Code tool. It is also available through and Meta AI for Mac. Meta's launch comes as real-time speech systems increasingly move beyond basic transcription, with speaker identification, multilingual conversations, and faster response times becoming key parts of AI-powered voice tools.
[4]
Meta launches Muse Voice Transcribe with real-time speaker detection and multilingual support
Meta is rolling out the model across Meta AI, Muse Code, its Model API and Meta AI for Mac. Meta has introduced Muse Voice Transcribe, a new artificial intelligence model designed to listen, transcribe and distinguish between speakers in real time. It can handle streaming transcription, identify more than 20 speakers in a conversation and detect when a person has finished speaking while the audio is being captured. Meta says that the model can also switch between languages mid-conversation and fine-tuned to recognise specific names, places and other terms. Built for speed and accuracy Meta stated that the model processes incoming audio in 80-millisecond slices and uses an adaptive delay system to determine when it has enough information to transcribe a word. The easier segments are processed almost instantly, while more complex phrases receive slightly more processing time. The company trained the system using reinforcement learning, balancing transcription accuracy with lower latency. Meta claims this helps the model improve the trade-off between speed and accuracy. The same architecture also handles speaker changes and detects conversational endpoints, allowing transcription and speaker tracking to work within a single system. Supports over 70 languages In a blog post, Meta stated that the model was trained on over 70 languages with 25 undergoing extensive validation. These include Chinese, French, Hindi, Japanese, Spanish and Vietnamese. Meta also showcased the model switching between English and Mandarin within the same sentence. The company also showcased its ability to handle long recordings involving multiple speakers without requiring manual cleanup. Meta sees the technology as a building block for more personalised AI assistants capable of understanding accents, interruptions, overlapping speech and multilingual conversations. Muse Voice Transcribe is rolling out immediately across several Meta products, including voice dictation in Meta AI and the company's Muse Code tool. The model will also be available through Meta's Model API and Meta AI for Mac.
Share
Copy Link
Meta Superintelligence Labs unveiled Muse Voice Transcribe, its first real-time audio perception model that combines live speech transcription, speaker separation, and endpointing in a single system. The AI model distinguishes between more than 20 speakers, handles code-switching across 70+ languages, and processes audio in 80-millisecond chunks using adaptive delay technology.
Meta Superintelligence Labs has launched Meta Muse Voice Transcribe, marking the company's entry into real-time audio transcription with an AI model that processes speech as it happens rather than after recording completion
1
. The real-time transcription system combines live speech transcription, speaker separation, and endpointing within a single architecture, eliminating the need for separate post-processing steps2
. Mark Zuckerberg demonstrated the model's capabilities on X, showcasing its ability to automatically distinguish between multiple speakers while seamlessly switching between languages mid-conversation1
.
Source: Gadgets 360
The real-time audio transcription model can identify and track more than 20 speakers in recordings longer than an hour, making it suitable for complex, multi-participant conversations
2
. Meta trained the system on over 70 languages, with 25 undergoing extensive validation for the initial release, including Chinese, French, Hindi, Japanese, Spanish, and Vietnamese3
. The model provides native support for five major Indian languages—Hindi, Tamil, Telugu, Kannada, and Malayalam—expanding accessibility across diverse linguistic markets2
. Real-time speaker detection happens as audio streams in, allowing the system to identify when someone starts or stops talking without delay4
.Muse Voice Transcribe processes incoming audio in 80-millisecond chunks, using adaptive delay to determine when sufficient information exists to transcribe each word accurately
1
. The AI model commits faster on simpler words while waiting slightly longer on complex phrases, balancing speed with accuracy through this intelligent processing approach4
. Meta employed reinforcement learning during training to optimize the trade-off between transcription accuracy and lower latency, helping the system make split-second decisions about when to finalize each transcribed token3
. This adaptive delay system allows the model to handle messy, real-world audio conditions while maintaining high accuracy across hour-long sessions1
.The model excels at handling code-switching, where speakers mix multiple languages within a single sentence—a common occurrence in multilingual conversations
1
. Users don't need to manually change language settings when speakers switch between languages, as the system automatically detects and adapts to linguistic shifts in real time2
. Meta demonstrated this capability by showing the model switching between English and Mandarin within the same sentence, maintaining accuracy throughout the transition3
. The system supports language, keyword, and context biasing, helping it identify words using information from both the audio stream and the broader conversational context2
.Related Stories
Meta has made Muse Voice Transcribe available through the Meta Model API at $3 per 1,000 audio minutes, which translates to approximately $0.18 per hour of transcription
2
. The model is immediately rolling out across several Meta products, including voice dictation in Meta AI, Muse Code, and Meta AI for Mac4
. Because the Meta AI for Mac app can power voice-enabled features across other applications, Muse Voice Transcribe will now handle dictation features on third-party services1
. Developers can access the model to integrate speech transcription capabilities into their own applications, with a demo version available on Meta's research blog1
.
Source: Engadget
Meta's release follows Google's launch of Gemini 3.5 Transcribe by less than a week, intensifying competition in the real-time audio transcription space
1
. While Google integrates its audio model into Android and Chrome, Meta's integration strategy remains focused on its own ecosystem and developer access through APIs1
. Meta positions Muse Voice Transcribe as a building block for personalized AI assistants capable of understanding accents, interruptions, overlapping speech, and multilingual conversations—capabilities increasingly demanded in modern voice interfaces3
. The model represents the latest release from Meta Superintelligence Labs, which has recently introduced its first dedicated coding agent, an open-weight model, and the Meta AI Mac app1
. As real-time speech systems evolve beyond basic transcription, speaker identification and faster response times are becoming critical differentiators in AI-powered voice tools3
.Summarized by
Navi
[3]
1
Technology

2
Policy and Regulation

3
Health