7 Sources
[1]
Global speech AI struggles to understand India: Report
A new national benchmark for speech recognition in India, 'Voice of India', has found a critical performance crisis for global AI models in the Indian market. While voice becomes the primary digital interface for millions in India, the benchmark reveals that leading global systems, including
[2]
India's homegrown AI revolution: How Sarvam AI outperformed global giants in key India-Centric tasks
Bengaluru-based Sarvam AI is redefining India's role in artificial intelligence by building foundational models that excel on tasks tailored for the nation's linguistic diversity. In recent evaluations, its OCR tool and Indic voice synthesis systems registered performance that beat well-known
[3]
Global Speech AI Struggles to Understand India: New National Benchmark 'Voice of India' Reveals
● Voice of India, a new national benchmark by Josh Talks & AI4Bharat evaluates leading speech-recognition systems across 15 languages and 35,000+speakers, revealing major performance gaps across demographics and real world speech ● Sarvam Audio claims top-tier rankings across major languages,
[4]
Sarvam AI Outshines Gemini and ChatGPT with 84.3% OCR Accuracy, Global Eyes on India
Sarvam AI Gains Global Backing as Vision Hits 93.28% Accuracy and Bulbul V3 Expands to 11 Indian Languages India takes a major step forward in artificial intelligence at a global scale. Sarvam AI, a Bengaluru-based startup, has surprised the tech community with AI models that perform better than
[5]
Saaras V3 explained: How 1 million hours of audio taught AI to speak "Hinglish"
The linguistic landscape of India is not a collection of neat, isolated boxes. It is a fluid, rhythmic blend where languages collide and merge in the middle of a single breath. For years, global speech recognition models have struggled with this reality, often tripping over the "Hinglish" or
[6]
Better than Google Gemini and ChatGPT? Indian startup Sarvam AI claims to beat global models
Launched ahead of the India-AI Impact Summit 2026, Bulbul V3 strengthens India's homegrown AI ecosystem, with real-time speech, enterprise features, and consent-based voice cloning. Bengaluru-based startup Sarvam AI has recently launched Bulbul V3, which is a new text-to-speech model designed for
[7]
Bulbul to Vision: Sarvam AI challenges global models with Indic stack
Indigenous models push India toward sovereign AI leadership If India's AI ambitions needed a pre-India AI Impact Summit flex, Sarvam AI delivered it loud and clear. Days before the India AI Impact Summit 2026 kicks off in New Delhi, the Bengaluru-based startup has rolled out a rapid-fire trio of
Share
Copy Link
A new national benchmark reveals that leading global AI systems from OpenAI and Microsoft struggle to understand how Indians actually speak. Sarvam AI, a Bengaluru-based startup, consistently ranks first across 15 Indian languages, achieving 93%+ accuracy while OpenAI's models trail by over 50 percentage points in the comprehensive evaluation.
A comprehensive national benchmark for speech recognition in India has revealed a striking performance crisis for global AI systems attempting to serve one of the world's largest voice-first markets. The Voice of India benchmark, developed by Josh Talks in collaboration with AI4Bharat at IIT Madras, evaluated leading Automatic Speech Recognition (ASR) systems across 15 languages and approximately 35,000 speakers, exposing significant limitations in how global AI models handle Indian languages
1
. The results challenge the readiness of voice-based AI for India's rapidly growing digital population, where voice is becoming the primary interface for millions of users.The benchmark results show that Bengaluru-based Sarvam AI consistently ranks first or second across almost every language and dialect tested, including major languages like Hindi and Bengali as well as regional ones like Odia and Assamese
3
. Sarvam Audio achieves 93%+ accuracy in critical regional dialects where global models falter. In stark contrast, OpenAI faces a massive performance disparity in Indian language transcription. While Google Gemini remains competitive with Sarvam, OpenAI's GPT-4o models trail by over 50 percentage points in accuracy compared to Sarvam in the overall average1
. Despite ChatGPT's global popularity, OpenAI's transcription models struggle immensely with Indian speech, registering over 55% Word Error Rate (WER). In languages like Maithili and Tamil, these models fail to transcribe nearly two out of every three words correctly3
.
Source: Digit
The Voice of India benchmark evaluates ASR performance using conversational speech collected from approximately 2,000 speakers per language, spanning a wide range of age groups, genders, regions, socio-economic backgrounds, device types, and acoustic environments
1
. Unlike many existing evaluations, it explicitly includes code-switched speech such as Hindi-English, Tamil-English, and Urdu-Hindi, as well as background noise and informal speaking styles common in everyday Indian conversations. The benchmark incorporates cluster-based geographic sampling across districts to capture how speech varies within a language's footprint, recognizing that pronunciation and vocabulary can shift significantly within 50-100 kilometers in India3
. Mitesh Khapra from AI4Bharat at IIT Madras emphasized that this represents "one of the most rigorous large-scale evaluations of speech recognition for Indian languages, containing district level cohorts with balanced representation across gender and age to truly reflect India's diversity"1
.The evaluation reveals that all models, including Sarvam, perform significantly better in Indo-Aryan languages like Hindi and Bengali at approximately 5-6% WER compared to Dravidian languages such as Tamil, Telugu, Malayalam, and Kannada at 15-20% WER
1
. Global speech systems often treat Hindi as a single, standardized language, but Hindi encompasses major dialects and accents such as Bhojpuri and Chhattisgarhi, each spoken by tens of millions of people. Bhojpuri alone has over 50 million speakers, a population larger than most European countries. Yet these dialects remain among the most challenging for AI systems, with even the best models seeing error rates jumping to 20-30% compared to sub-10% in standard Hindi3
. Despite Urdu being linguistically similar to Hindi, OpenAI models perform poorly in Urdu with 35.4% WER, while Sarvam Audio maintains high accuracy at 6.95% WER1
.Founded in 2023 by Dr. Vivek Raghavan and Dr. Pratyush Kumar, Sarvam AI set out to create compact, efficient foundational models capable of running on phones and modest infrastructure while effectively handling India's complex linguistic landscape
2
. The company's Saaras V3 model was trained on over one million hours of multilingual audio data, capturing the raw reality of Indian speech across various accents, background noise levels, and acoustic conditions5
. This massive training scale allows the model to handle code-mixing as a primary feature rather than treating it as noise. Saaras V3 achieves a Word Error Rate of 19.3% on the IndicVoices benchmark, consistently outperforming frontier models like GPT-4o and Gemini 3 Pro when tested in India5
. The model utilizes a streaming-first architecture with causal attention, delivering a time-to-first-token of under 150 milliseconds for real-time voice applications5
.
Source: Digit
Related Stories
Sarvam AI's Vision tool, an optical character recognition model designed for native Indian scripts, registered higher OCR accuracy than widely used global models on benchmarks for Indian language document recognition
2
. Reports indicate the Vision model achieved 84.3% accuracy, with some configurations reaching 93.28% accuracy4
. The company's Bulbul V3 model for voice synthesis generates expressive text-to-speech output across 11 Indian languages. Independent tests showed that Bulbul V3 handled numerals, named entities, and code-mixed text more effectively than several competitive systems2
. These AI models for India demonstrate that tailored engineering and careful data curation can deliver strong results for complex localized problems that large generic systems sometimes overlook.
Source: Analytics Insight
Sarvam AI's approach aligns with growing interest in sovereign AI solutions built within the country and designed to meet local regulatory and privacy expectations
2
. By focusing on India's unique challenges, this philosophy contrasts with dominant global AI narratives that prioritize breadth of capability over local specificity. Tools that reliably recognize text across diverse document layouts and languages can streamline workflows in banking, education, and public services where paper-based and multilingual communication remains common. Voice technologies that understand India's vernacular languages can broaden digital service reach, especially in regions where English is not predominant. Meanwhile, Microsoft STT is not supported for nearly half the languages tested, including major regional languages like Punjabi, Odia, and Kannada3
. Meta's massive 7B parameter model is only approximately 4% more accurate than its much smaller 1B parameter model on average across Indian languages, highlighting efficiency gaps in global approaches1
. As India positions itself as a serious AI innovator, the success of Indian AI in handling Hinglish and other code-mixed languages suggests that understanding local context may be as critical as computational scale in building effective AI systems for diverse markets.Summarized by
Navi
[2]
[3]
[4]
1
Science and Research

2
Policy and Regulation

3
Technology