Regional leaderboards ignore Global South languages
For years, the assumption has been that better data would fix AI bias. Feed more Hindi, Swahili, or Arabic text into a model, the logic went, and it would finally serve the people who speak those languages. A new position paper from researchers at the intersection of AI governance and global development turns that assumption on its head. The problem, they argue, is not missing datasets. It is missing institutions.
The paper, published on the arXiv preprint server and drawing on a consultation with 58 AI practitioners in India, makes a startling claim: high-quality regional benchmarks already exist. IndicSUPERB, MILU, and LAHAJA cover Indian languages with rigor. IrokoBench does the same for African languages. AlGhafa serves Arabic speakers. These are not amateur efforts or afterthoughts. They are professionally constructed evaluation tools that could tell us, with reasonable confidence, whether a model actually understands Tamil or Yoruba or Swahili.
Yet none of these benchmarks appear on the global leaderboards that tech companies, researchers, and journalists treat as the definitive word on AI capability. The result is a perverse feedback loop. A model that fails spectacularly on Hindi comprehension can still rank near the top of a global leaderboard, because that leaderboard simply does not ask the question. The failure is not invisible to everyone — the practitioners who built those benchmarks have documented it thoroughly — but it is invisible to the people who make decisions.
This is where the human role disappears. Once upon a time, a linguist or a regional expert would have been consulted to judge whether a model handled a language correctly. That judgment was slow, expensive, and hard to scale, but it existed. Today, the leaderboard has replaced the expert. Its numbers are treated as objective truth, even though the choices behind those numbers — which benchmarks to include, which languages to test, which tasks to measure — are made by a small group of institutions with commercial interests.
The paper’s authors are careful to note that this is not a conspiracy. No one deliberately excluded Indian or African benchmarks from global leaderboards. The exclusion is structural, not intentional. The exclusion is structural. Leaderboards are run by organizations that answer to funders and shareholders. Those organizations respond to pressure from paying customers, and their customers are overwhelmingly in the Global North. When a European bank complains that a model fails on German legal text, the problem gets fixed quickly. When a Hindi-speaking farmer in Uttar Pradesh cannot get a reliable answer from the same model, no one with power hears about it.
The asymmetry is the point. Commercial pressure corrects leaderboard failures only when the people affected have economic leverage. The Global South, for all its population size, lacks that leverage in the AI marketplace. So failures affecting Hindi, Swahili, or Arabic speakers persist indefinitely — documented, acknowledged in academic papers, but never addressed in practice. The gap between what the leaderboard claims and what the model actually does becomes a permanent feature of the landscape.

India makes an especially sharp case study because it has everything except the institutional mechanism. 1.4 billion people. 22 scheduled languages. High-quality benchmarks built by local researchers who understand the linguistic nuances that global evaluators miss. What India lacks is a trusted aggregation point — a place where all those benchmarks come together, where results are comparable, and where the findings carry enough authority to force change.
The consultation with 58 AI practitioners revealed a striking consensus. These are people who work with AI daily, who understand its technical capabilities and limits. They did not ask for more data or better models. They asked for governance. They asked for independent oversight. They asked for disclosure-based conflict management, so that the organizations running leaderboards would have to reveal whose interests they serve.
This preference for institutions over technology is itself a kind of revelation. The AI discourse has been dominated by the idea that scale solves everything — more parameters, more training data, more compute. The practitioners in this study are saying the opposite. The bottleneck is not technical. It is political. It is about who gets to decide what counts as good performance, and whose failures are allowed to matter.
The paper proposes a concrete alternative: regional leaderboards with independent governance from the start. Not a global board that grudgingly adds a few regional benchmarks, but a separate infrastructure built by and for the regions it serves. Such a board would have its own conflict-of-interest policies, its own mechanisms for updating metrics as languages evolve, its own accountability to the communities it represents.
The proposal sounds modest, but its implications are not. If regional leaderboards gain authority, they would change the incentive structure for AI companies. A model cannot simply dominate a global leaderboard and claim victory if it fails regional evaluations that carry weight. The companies would have to care about performance in Hindi and Swahili and Arabic, not because they are ethically compelled, but because their reputation would depend on it. This is the mechanism that makes the proposal viable: it aligns corporate self-interest with linguistic justice.
There is a deeper irony here that the paper does not dwell on but that lingers in the background. The AI systems that are supposed to be so intelligent cannot even be properly evaluated without human institutions to judge them. The leaderboard, which was supposed to remove subjective human judgment from the assessment process, turns out to be the most subjective part of all. It just hides its subjectivity behind numbers.
The researchers are not naive about the difficulty of their proposal. Building an independent governance body is slow, unglamorous work. It does not generate headlines or attract venture capital. It requires sustained commitment from people who could probably earn more money doing something else. But the alternative, they suggest, is a world where the most powerful technology ever created simply does not work for most of humanity, and no one is even required to notice.

The puzzle at the heart of this paper is that the solution is so obvious once stated. Of course evaluation should be independent. Of course the people whose languages are being tested should have a say in how they are tested. Of course an institution that claims to measure something should be accountable to the people it measures. These are not radical ideas. They are the basic principles of any functioning regulatory system, applied to a domain that has so far avoided regulation entirely.
What makes the paper uncomfortable is what its findings imply about the AI industry’s self-image. The industry likes to see itself as meritocratic, driven by data and evidence rather than politics and power. The leaderboard is the symbol of that self-image — a clean, objective ranking that anyone can verify. This paper suggests that the symbol is hollow. The ranking reflects choices, and the choices reflect interests, and the interests are concentrated in a few wealthy regions of the world.
The practitioners who participated in the consultation seem to have understood this intuitively. They did not need a study to tell them that leaderboards were failing them. They had been living with that failure for years. What the study provided was a framework for talking about it — a way to name the problem and point toward a solution without descending into conspiracy theories or despair.
The paper’s final insight is that progress and consequence rarely share the same time horizon. The benchmarks exist. The failures are documented. The solutions are proposed. But institutions take time to build, and meanwhile the models keep getting better at the things the leaderboards measure, and the gap between what the leaderboards measure and what people actually need keeps growing. The window for action is now, while the infrastructure is still being shaped. Once the institutional patterns are set, they will be as hard to change as any other entrenched system.
For the 1.4 billion people in India, for the billions more across the Global South, the stakes could not be higher. They are being asked to trust a technology that has never been properly tested on them, judged by standards they had no role in setting. The paper does not ask for charity or special treatment. It asks for something more basic: that the measure be fair. That the people who are being measured get a seat at the table. That the leaderboard, which has quietly replaced human judgment as the arbiter of AI quality, be held to the same standards of accountability that we expect from any other institution that claims to represent reality.
Sources
1. MILU
2. LAHAJA
