2 Sources
[1]
Voice AI Models: Voice AI models struggle to understand speech, creating challenges for robot training: Humyn Labs
Voice AI models frequently encounter hurdles when it comes to grasping the intricacies of human speech. Variances in dialects and overlapping conversations can lead to significant transcription inaccuracies. Furthermore, extended silences may disrupt the models' ability to retain contextual
[2]
Humyn Labs Benchmarks Voice AI: Exposes Critical Failures in Human-Robot Interaction
BRIDGE evaluates how leading AI models like Sarvam v3, Gemini 3 Pro, ElevenLabs among others perform across the world languages for humans to interact with robots effectively Humyn Labs today launched the second edition of BRIDGE, a benchmark report that reveals the gap between human speech and AI
Share
Copy Link
Humyn Labs released its second BRIDGE benchmark report exposing critical gaps in how Voice AI models understand human speech. The study tested 23 models across 23 languages and found overlapping speech alone increased transcription errors from 41.2% to 45.2%. Regional dialects worsened performance further, with Bengali showing a 9-point error gap between standard and regional variants, raising concerns for collaborative robots and human-robot interaction.

Voice AI models continue to struggle with understanding human speech in real-world environments, according to the second edition of the BRIDGE benchmark report released by physical AI research lab Humyn Labs
1
2
. The comprehensive study evaluated 23 voice AI models across 23 languages, revealing that overlapping speech, regional dialects, long pauses and code-switching create significant transcription errors that undermine human-robot interaction. These accuracy gaps matter because voice is emerging as a critical interface between humans and collaborative robots, and failures in AI voice processing technology directly impact automation, customer experience and trust in real-world settings.The BRIDGE benchmark report found that overlapping speech alone pushed the average error rate from 41.2% to 45.2%
2
. Regional dialects compounded the problem dramatically. Bengali scored 42.4% in its standard form but jumped to 51.0% in a regional dialect outside Kolkata, a gap of nearly 9 percentage points caused by dialect variation alone. The dialect challenge extends beyond Indic languages: Argentinian Spanish recorded a 7.85% word error rate compared to 16.04% for Venezuelan Spanish, more than doubling across dialects of the same language. These findings confirm that Voice AI models fail to understand human speech consistently across linguistic variations, creating barriers for Physical AI systems that need to operate across diverse markets and populations.Model choice emerged as one of the largest levers affecting real-world accuracy. ElevenLabs, the top performer across five non-Indic languages, averaged 5.8% error compared to 24.6% for the widely used GPT-4o-mini-transcribe—a gap of more than 4x on identical audio
2
. The study tested models including Sarvam v3, Gemini 3 Pro, Soniox, Speechmatics and Gnani Vachana across more than 200 hours of human-verified real-world audio collected across two to three districts per language. With 5.5 billion people speaking languages other than English, provider selection becomes critical for enterprises deploying voice-enabled automation. Manish Agarwal, Co-Founder at Humyn Labs, emphasized that "voice accuracy becomes a business imperative, not just a technical metric" for Physical AI builders and enterprises evaluating whether models are ready to scale across markets2
.Extended silences significantly disrupted Voice AI models' ability to retain contextual relevance. Brazilian Portuguese calls showed an 18.8% error rate for gaps exceeding 150 seconds versus 12.4% for shorter gaps of around 35 seconds
2
. This pattern replicated internationally, confirming that models losing the thread across long pauses is not limited to specific language families. The inability to maintain context through natural conversational pauses undermines performance in real-world AI settings where humans and collaborative robots must work together effectively. These transcription errors can diminish user satisfaction and impact the reliability of automation systems that depend on accurate voice commands and responses.The BRIDGE benchmark revealed that Voice AI models exhibit three distinct failure modes at similar overall error scores. Most models—19 out of 23 tested—mishear words and substitute the wrong ones through substitutions
2
. A smaller group, including OpenAI's transcribe models, Speechmatics and Gnani Vachana, fail primarily through omissions, dropping words rather than mistranscribing them, with deletions accounting for 38-39% of their total errors. Gemini Flash demonstrated a third failure pattern through fabrications, introducing words that were never spoken and adding invented content equal to 9.3% of the reference transcript's length. The distinction matters commercially because workflows that can absorb a missing word may not tolerate fabricated content, particularly in safety-critical human-robot interaction scenarios.Related Stories
BRIDGE separates script choice from genuine transcription errors, a distinction standard scoring misses. In Bengali, Soniox and Sarvam v3 recorded raw error rates of around 20% to 21%, roughly half of which came from writing English loanwords in a different script rather than from mishearing them
2
. Gemini 3 Pro recorded 9.2%, of which only 1.6 percentage points were script mismatch. This granular analysis helps enterprises understand whether errors stem from actual misunderstanding of speech or from formatting decisions, informing better model selection for multilingual deployments where code-switching between languages is common.The benchmark tested the economic trade-offs of model routing strategies. The best single model reached 10.7% loanword-adjusted error, while selecting the best model per language improved performance to 9.7%, and a theoretical best model per call reached 8.9%
2
. One model won 78.6% of files outright, suggesting that hybrid routing strategies could balance cost and accuracy. Ishank Gupta, Co-Founder at Humyn Labs, noted that "Physical AI cannot learn the real world through vision alone" and that sound carries information about people, actions, distance, environment and intent that robots operating alongside humans need to interpret reliably. BRIDGE provides the evaluation layer testing speech models across conditions that understand human speech in noisy, conversational density environments where collaborative robots must function effectively alongside people.Summarized by
Navi
09 Feb 2026•Technology

02 Jan 2026•Technology

08 May 2026•Technology

1
Technology

2
Technology

3
Science and Research
