2 Sources
[1]
AI4Bharat collects ten trillion tokens of data to power AI in Indian languages
"Several startups, academic institutes and deeptech institutes are using this data to build their own models to accelerate the adoption of language technologies" said Mitesh Khapra, cofounder of AI4Bharat.Chennai-based AI4Bharat is collecting ten trillion tokens of language data from everyday
[2]
AI4Bharat Collecting 10 Tn Tokens To Build Next Generation Of AI
AI4Bharat claims to have sourced the data from voice samples of users across several demographics and professions IIT Madras-incubated artificial intelligence (AI) lab, AI4Bharat, is reportedly collecting 10 Tn tokens of language data to build the "next generation of AI services". For context,
Share
Copy Link
AI4Bharat is collecting 10 trillion tokens of language data from across India to develop AI models that can effectively understand and process Indian languages, aiming to bridge the gap in AI accessibility for the country's linguistically diverse population.
AI4Bharat, an IIT Madras-incubated artificial intelligence lab, has embarked on a groundbreaking project to collect ten trillion tokens of language data from across India. This massive undertaking aims to power the next generation of AI services tailored for Indian languages
1
. The initiative, known as the "Ten Trillion Token" project, seeks to address the unique challenges posed by India's linguistic diversity and create AI models that can effectively understand and process Indian languages.Over the past three years, AI4Bharat has conducted an extensive data collection campaign, covering almost every district in the country and encompassing all 22 official languages of India. The collected data includes:
Mitesh Khapra, co-founder of AI4Bharat, emphasized the importance of this diverse dataset: "We have ensured that we collect voice samples split across several demographics, across different professions, blue collar and white collar"
1
.The collected data is expected to have wide-ranging applications, including:
These use cases demonstrate the potential impact of language-aware AI on various sectors of Indian society
2
.AI4Bharat has adopted an open-source approach to accelerate the development and adoption of language technologies. Khapra stated, "Our data, models and scripts are open sourced. You can build on top of that"
1
. This approach has enabled various stakeholders, including startups, academic institutions, and deep tech companies, to utilize the collected data for building their own models.Related Stories
The Ten Trillion Token project aims to address a critical gap in current AI technologies. While English-language data is abundant on the internet, making it easy to train AI models, the same is not true for Indian languages. Each of India's 22 major languages has its own script, grammar rules, and cultural context, presenting unique challenges for AI development
1
.The successful completion of this project could have far-reaching implications for AI accessibility in India. It could enable:
By building native Indic models that support Indian languages "not as an afterthought," AI4Bharat aims to create AI systems that truly work for India's diverse population
2
.Summarized by
Navi
03 Jun 2025•Technology

09 Feb 2026•Technology

24 Sept 2025•Science and Research

1
Technology

2
Policy and Regulation

3
Technology
