Gates Foundation launches coalition of 60 organizations to build language datasets for AI accessibility

5 Sources

Share

The Gates Foundation has convened Anthropic, Google, OpenAI, and 57 other organizations to make AI tools accessible in underrepresented languages. The coalition aims to reach over 3 billion people in five years by building more representative language datasets, addressing global inequality in AI technology.

Gates Foundation Launches Coalition to Build Language Datasets for AI

The Gates Foundation announced Monday the formation of a coalition of 60 organizations dedicated to making AI tools accessible in underrepresented languages

1

. The partnership brings together Anthropic, Google, and OpenAI Foundation alongside frontier AI labs, corporations, and philanthropies already working to expand language availability in AI systems. The coalition aims to reach more than 3 billion people over five years by better coordinating existing efforts to build more representative language datasets

2

.

Gates Foundation CEO Mark Suzman emphasized that this work must continue "full speed ahead," even as some major AI companies call for slowing down development of advanced models

3

. "Even if AI was frozen right now—which I don't expect and I am not calling for—we would want to be building these language sets and making them usable with the tools that we have available right now," Suzman told The Associated Press

1

. He argued that governments must regulate impacts on cybersecurity and children's development while extending AI's humanitarian applications to poor communities currently "shut out" from the technology.

Source: AP

Source: AP

Addressing Global Inequality Through Language Accessibility

The initiative follows last week's Goalkeepers report, where the foundation committed $1 billion toward AI-focused efforts to improve health outcomes, upgrade educational tools, and inform small farmers' practices around the world

4

. Unrepresentative language data creates serious consequences for global inequality. The report warned that a model could mistranslate a pregnant Malawi woman saying her "water has broken" into the direct English that she'd "thrown away water," demonstrating how cultural and linguistic diversity failures can impact critical health, education, and agriculture applications

5

.

Mozilla Data Collective CEO E.M. Lewis-Jong, whose company was incubated by the Mozilla Foundation, called this the "original sin" of AI development

1

. Many AI tools were trained on language data scraped from the internet without ethical data collection practices. "The internet is not a representative space," she said. "Why would you think that you were going to get a culturally diverse and representative system out of something that was predominantly trained on Reddit?"

2

Her data-sharing platform seeks to help communities upload cultural and linguistic data sets on their own terms rather than having that data taken from the web without explicit consent.

Major Tech Companies Commit to Language Expansion

Google has been funding Project Vaani, an effort to collect more than 150,000 hours of audio across every district in India

3

. According to Google Senior Vice President James Manyika, the initiative underscores the need to gather speech data on dialects within languages. To accomplish this, Google is working with local partners to record speech in the field, recognizing that true AI accessibility requires capturing linguistic variations within broader language categories

4

.

Anthropic, whose CEO recently published an essay calling for industrywide cooperation on decelerating advancements, was already working with the Gates Foundation to speed up vaccine developments and bolster its chatbot's data set of local crops

5

. Elizabeth Kelly, head of beneficial deployments at Anthropic, acknowledged the company's products "lag in many African languages in particular." She emphasized the urgency: "We're acutely aware that we can't achieve any of the benefits we want to see in terms of improving patient outcomes or improving literacy and numeracy for students across the globe unless we actually get this language piece right"

1

.

Coalition Structure and Future Coordination

Details regarding AI governance and the coalition's operational structure are still being finalized, according to Mark Suzman

2

. A secretariat will track each signatory's commitments, and the Gates Foundation may nudge partners to fill larger gaps when necessary. This coordination mechanism aims to prevent duplication of efforts while ensuring comprehensive coverage of underrepresented languages across different regions and use cases.

The collaboration represents a significant shift in how frontier AI labs approach language datasets for AI development. Rather than continuing to rely on internet-scraped data that perpetuates existing biases, the coalition of 60 organizations is working toward systematic, community-driven data collection that respects cultural contexts. Watch for announcements about specific language targets, governance frameworks, and measurable milestones as the coalition operationalizes its five-year plan to transform AI accessibility for billions currently excluded from these technologies.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved