
Image: Olkeri
By Olkeri.space
The AI Language Gap: Why Most of the World Gets Worse AI
Models trained mainly on English serve billions of people poorly. Here is what that costs and who is fixing it.
Read this story in: Français · Deutsch · Español
Artificial intelligence works best in English, reasonably well in a handful of major languages, and poorly or not at all in most of the thousands of languages people actually speak. That gap determines who benefits from the technology.
Why the gap exists:
Language models learn from text. The volume of digital text available differs by orders of magnitude between languages: English dominates, followed by a few widely digitised languages, while most of the world's languages have very little written material online.
A model trained predominantly on English text develops strong English capability and weaker capability elsewhere, roughly in proportion to the data available. For languages with almost no digital corpus, performance can be unusable.
The problem compounds. Poor performance means fewer people use these systems in their language, which generates less new text, which keeps the data scarce.
Who is affected:
The severity varies. Major languages including Chinese, Spanish, Arabic, Hindi, Portuguese, French, Russian and Japanese have substantial resources and reasonable model performance, though still below English.
Below that sit hundreds of languages with tens of millions of speakers each and thin digital resources: Swahili, Amharic, Yoruba, Hausa, Igbo, Bengali dialects, Javanese, Tagalog, Vietnamese regional varieties, Kazakh, and many more.
Below that again are thousands of languages with limited written tradition or small speaker populations, for which no commercial case for model support exists at all.
Africa is the most affected continent, with over 2,000 languages and minimal representation in training data. Papua New Guinea alone has more than 800 languages.
What it costs:
The consequences are concrete. A farmer who speaks only a local language cannot use an agricultural advisory chatbot. A patient cannot describe symptoms to a diagnostic tool. A citizen cannot access a government service. Voice interfaces, which matter most where literacy is limited, are precisely where performance gaps are widest.
The result is that the populations who would benefit most from AI-mediated access to information and services are those the technology serves worst.
There is also a cultural dimension. Systems trained on English text encode assumptions, references and framings from English-speaking contexts, and answers reflect them even when translated.
Who is working on it:
The most significant work is community-driven. Volunteer and academic networks across Africa have built datasets, benchmarks and models for African languages, often with minimal funding, and their work is now the foundation most African language technology builds on.
National programmes exist too. Iceland funds Icelandic language technology explicitly as cultural policy. India has invested substantially in Indian language datasets and models given that most of its population does not use English. Sweden, Finland, Estonia, Latvia and others fund their own language models. Gulf states fund Arabic capability. Japan, South Korea and Taiwan fund their own.
The economics differ fundamentally between these cases: large-market languages attract commercial investment, small languages require public funding, and languages that are both small and poor get neither.
What would help:
The practical needs are unglamorous: digitised text, transcribed speech, evaluation benchmarks, and speakers paid properly to create and validate data.
Open-weight models help disproportionately here, because a community with a modest dataset can adapt an existing capable model rather than train one from scratch, which is otherwise financially impossible.
The gap will not close on its own. It closes where someone decides a language is worth the investment, and that decision is currently made by markets for most of the world.