88 percent of AI-generated stories set in Indian contexts contain cultural inaccuracies. That single number comes from research presented at the 43rd International Conference on Machine Learning. It captures something fundamental about how the world’s most powerful AI systems are tested and about who gets left out of the picture entirely.
Every time a government procures AI software, somebody checks an AI leaderboard. These rankings tell the world which models perform best. They shape procurement decisions worth billions of dollars. They direct investment capital. They guide the research priorities of thousands of engineers. One study found that companies spend hundreds of thousands of dollars in computing resources just to improve their standing on these lists.
When a leaderboard declares a model the best, that signal travels fast. The model lands on procurement shortlists. Investment memos quote the score. Contracts follow. The rankings carry real economic and social weight, far past what their creators originally intended.
The trouble is that these rankings were mostly designed for English speakers in wealthy Western countries. Hindi, Swahili, Arabic, Bengali, and hundreds of other languages spoken by the majority of humanity are largely absent from the tests. A model can sit at the very top of global rankings while failing catastrophically for a farmer in rural India, a health worker in Nigeria, or a student in Indonesia. The score says “best.” The deployed product says something else entirely.
A position paper published at ICML 2026, by researchers from IIT Kharagpur, Shunya Labs, and Nasscom, argues that this failure isn’t accidental. High-quality benchmarks for Indian, African, and Arabic languages already exist. Researchers have built them. They are rigorous and peer-reviewed. Global AI leaderboards simply don’t use them. The paper argues that the core problem is harder to fix than missing data. It’s institutional. These leaderboards have no independent oversight, no formal conflict-of-interest policies, and no mechanisms that require them to update as knowledge grows. Because the Global South has no commercial pull over these systems unlike enterprise clients in wealthy nations the failures persist year after year, documented but unaddressed.
“The barrier is not missing data,” the paper states plainly. “The barrier is institutional design.”
The AI Leaderboard Problem: Rankings That Rule the World
AI leaderboards started as simple academic scoreboards. Researchers wanted an easy way to compare how models performed on standard tasks. Over the years, these rankings became something far more consequential.
Governments now cite them in procurement guidance. The US Office of Management and Budget referenced them in federal AI acquisition policy. Enterprises use them to select vendors. Investors quote benchmark scores in due diligence documents. Companies redirect engineering teams to chase higher positions.
The paper’s authors use a striking phrase to describe what leaderboards actually do in practice: they function as “implicit loss functions.” In plain language, what a leaderboard rewards is what engineers optimize their systems to achieve. If a ranking tests AI models on English text accuracy, engineers build toward English text accuracy even when the model will ultimately serve speakers in Mumbai, Lagos, or Jakarta.
That optimization pressure flows downward through every decision in the development process. It shapes what data gets collected, what tasks get prioritized, and which user populations receive attention. The failures that result are systematic, not coincidental.
The Languages Left Behind by AI Benchmarks
One concrete example illustrates the scale of the problem. The HuggingFace Open ASR Leaderboard one of the most referenced rankings for speech recognition AI includes a section called “Multilingual ASR Evaluation.” It covers five languages: German, French, Italian, Spanish, and Portuguese. Every single one is European.
Hindi has over 600 million speakers. Arabic has more than 370 million. Bengali has 270 million. Indonesian has 200 million. Swahili has 100 million. None of them appear on the list.
The bias shows up across the most commonly used AI tests. Research cited in the paper found that 84.9 percent of geography questions in MMLU one of the most popular benchmarks for measuring AI “intelligence” focus exclusively on North American or European regions. Just 15 percent has been adapted for non-Western contexts.
The training data tells the same story. English makes up 43.8 percent of Common Crawl, a massive dataset used to train major AI systems. English is a first language for less than 20 percent of the world’s population. Arabic, the fifth most widely spoken language on Earth, accounts for less than 1 percent of that training data. Over 2,000 African languages are, for practical purposes, invisible to the AI systems now being deployed across the continent.
The notable thing is that better benchmarks already exist.
Indian researchers have built IndicSUPERB, MILU, and LAHAJA to test AI performance across Indian languages.
IrokoBench covers 16 African languages.
AlGhafa targets Arabic specifically.
SEA-HELM covers Southeast Asian languages.
These are serious, peer-reviewed tools built by accomplished research teams.
Global AI leaderboards simply don’t use them. No technical barrier prevents their inclusion. The paper argues that no governance structure requires it so it doesn’t happen. The benchmarks sit there, available and ignored.
When the Numbers Break Down: Documented Failures Across Industries
The consequences of poorly calibrated AI leaderboards are not theoretical. They appear in hospitals, on farms, and in government offices with measurable, documented frequency.
In language AI, the performance gaps are stark. GPT-4 produces significantly more hallucinations false information presented confidently as fact when working in Hindi compared to English. The LAHAJA benchmark found a 15 to 30 percent drop in speech recognition accuracy across regional Hindi varieties. The best-performing model on the IndicParam benchmark, Gemini-2.5, achieved only 58 percent accuracy on low-resource Indic languages. GPT-4 scored just 45 percent figures that would be considered catastrophic failures if they appeared on English-language evaluations.
Agentic AI systems, programs that take autonomous actions on a user’s behalf, drop from 60 percent performance under controlled benchmark conditions to just 25 percent in real production settings. That’s a collapse of more than half the apparent capability, hidden by tests that don’t reflect actual conditions.
In medical AI, the gaps carry genuinely serious consequences:
A 2025 study found that vision-language models analyzing chest X-rays consistently underdiagnose marginalized groups, with the worst error rates for Black female patients.
Western cardiovascular risk models widely used as clinical reference tools misclassified nearly 80 percent of 4,975 first-time heart attack patients in India as low or moderate risk. These patients were experiencing acute cardiac events. The models had never been re-validated against South Asian physiology or lipid profiles before adoption as clinical standards.
In agriculture, the numbers are equally sobering:
The PlantVillage crop disease detection system reported 99.35 percent accuracy on its North American laboratory benchmark.
When deployed in Tanzania to detect cassava disease affecting smallholder farmers, accuracy fell to 49 percent. That’s not a performance drop it’s a coin flip for farmers whose livelihoods depend on early detection.
In weather forecasting, a final irony emerges. Africa operates at roughly one-eighth of the WMO’s recommended surface observation station density. Major AI weather models trained on that sparse global data inherit those gaps. The communities most exposed to weather stress receive the least reliable AI predictions.
In every case the paper examines, a technical solution is known and available. Better local data, regional fine-tuning, validation against deployment populations the tools exist. The obstacle is institutional. No governance structure requires validation on affected populations before deployment. No appeals process gives those populations any recourse after the fact.
The Governance Gap in AI Leaderboards
To understand why these failures persist, it helps to understand what most AI leaderboards currently lack entirely.
A conflict of interest, as the paper precisely defines it, exists when the same parties who benefit from high rankings also control the evaluation process. This isn’t a claim about bad faith. It’s an observation about structural incentives that can bend outcomes even when everyone involved is trying to act properly.
The paper uses the HuggingFace Open ASR Leaderboard as an illustration. Some co-authors of that leaderboard also developed models that rank at the top of it. Related training datasets came from overlapping research teams. Core architectural components were built by people with ties to the same leaderboard. No published conflict-of-interest policy governs any of these overlaps. No formal mechanism lets developers challenge rankings they believe are unfair.
This pattern is not unique to one organization. A study cited in the paper found that a single AI provider evaluated 27 model variants privately before public release. The top two providers combined received 39.6 percent of all arena evaluation data. Meanwhile, 83 open-weight models together received only 29.7 percent. Academic peer review requires referees to declare conflicts and step back from papers where they have stakes. AI leaderboards have no equivalent structure.
Transparency is getting worse, not better. The Foundation Model Transparency Index found that average scores across major AI organizations fell from 58 in 2024 to 40 in 2025. Companies are most secretive about training data, computing resources, and how their models behave after deployment.
There’s also what the paper describes as Goodhart’s Law in action. The principle states: when a measure becomes a target, it stops being a good measure. AI companies pour resources into benchmark scores. The scores gradually lose their connection to genuine real-world capability. English speakers may experience genuine improvements before the metrics break down. People whose needs were never captured in the metrics receive only the costs gaming, contaminated test sets, misallocated engineering effort without ever seeing the benefits.
Why Markets Won’t Solve the AI Leaderboard Problem
One reasonable objection to the paper’s argument is worth examining: won’t market pressure eventually force AI companies to serve the Global South better? If systems fail in India or Nigeria, surely they’ll eventually lose enough business to care?
The paper draws a pointed analogy to global health research funding. The “10/90 gap” is a documented phenomenon in which less than 10 percent of health research addresses conditions causing 90 percent of the world’s disease burden. Cancer research attracts enormous investment because patients in wealthy nations can pay for treatment. Malaria research remains underfunded because most affected populations cannot.
AI evaluation infrastructure follows the same structural logic. When a speech recognition system fails for enterprise clients in the United States, those clients escalate the problem, threaten contract cancellations, and demand fixes. Engineering resources get redirected. When the same system fails at higher rates for Hindi or Hausa speakers, the problem gets logged in release notes as a “scope limitation.” Documented. Not fixed.
The paper is blunt about the human cost: “A farmer in rural India interacting with a government voice assistant, or a health worker in Nigeria using an AI diagnostic tool, consumes AI systems selected through leaderboard-influenced procurement, yet has no seat at governance tables.”
Research on what scholars call “algorithmic transference” adds another dimension. When AI fails, people don’t merely stop trusting the technology. They stop trusting the institutions that deployed it. Governments and enterprises suffer credibility damage alongside the AI companies. The failure ripples through social trust in ways that compound over time.
The paper identifies four specific mechanisms through which Global South populations face a steeper disadvantage. There’s escape route asymmetry: a procurement team in Seattle can commission independent evaluations if leaderboards seem unreliable; most teams in the Global South cannot. There’s bug-fix priority asymmetry: failures for enterprise clients in wealthy countries draw engineering responses; failures for lower-resource language speakers get characterized as expected variance. There’s institutional redundancy: the Global North has overlapping checks on AI quality academic review, enterprise testing, regulatory bodies, consumer advocates while many Global South countries have only the leaderboard as a quality signal. Finally, there’s optimization lock-in: as AI models train on deployment data from English-speaking markets, the performance gap for other languages widens with each new model generation.
India: A $1.2 Billion Problem Without the Right Infrastructure
India makes for an instructive case study. At 1.4 billion people, with 22 constitutionally recognized languages and more than 80 regional varieties of Hindi alone, the country presents extraordinary AI deployment complexity.
India’s government is investing seriously. The IndiaAI Mission carries a $1.2 billion budget one of the world’s largest national AI programs. That mission references global benchmarks in its policy documents. Those benchmarks don’t adequately represent Indian linguistic reality.
The paper identifies a fundamental asymmetry: India has the technical tools but lacks the institutional infrastructure. Research groups at IITs and IISc have built rigorous, respected benchmarks. IndicSUPERB tests speech processing across multiple Indian languages under clean, noisy, and telephone-call conditions. LAHAJA evaluates regional Hindi varieties specifically. MILU tests language understanding across eight knowledge domains with Indian-specific content. Svarah measures how well AI handles Indian English speech patterns and the constant code-mixing the natural switching between languages mid-sentence that characterizes everyday speech across much of the country. IISc-MILE provides evaluation for the complex morphology of Tamil and Kannada, where words are built from chains of suffixes that standard word-error metrics completely mishandle.
These tools are credible and peer-reviewed. No trusted, independent body aggregates them into a unified ranked evaluation.
AI4Bharat, a research group at IIT Madras, launched the Indic LLM Arena in November 2025. The paper views this as a genuine step forward. It also uses AI4Bharat’s structure to illustrate the governance challenge: that organization creates benchmarks, builds models, and operates the leaderboard, all under one roof. The paper doesn’t describe this as improper. It points out that concentrated expertise is a natural feature of early-stage ecosystems. The question is whether governance structures get built before that concentration becomes difficult to manage transparently.
What India’s AI Community Wants: Survey Findings
Between December 2025 and March 2026, Nasscom surveyed 82 AI practitioners about what governance structure they’d want for a regional Indian AI leaderboard. The results were clear on the most basic question.
Zero respondents chose to leave governance informal. Every single participant wanted some form of formal governance structure in place.
Within that unanimous demand, preferences revealed a specific direction:
64 percent wanted non-government stewardship either a Nasscom-led body (32.9 percent) or an independent non-profit (31.7 percent)
21 percent favored government-driven governance through bodies like MeitY or IndiaAI
9.8 percent preferred an academic consortium
68 percent wanted conflicts handled through disclosure and recusal rather than pre-emptive exclusion of parties with potential conflicts
76 percent preferred hybrid evaluation combining AI judges and human reviewers
74 percent wanted quarterly submission windows four opportunities per year for models to be evaluated
The community rated “becoming a credible standard” as their top success metric, scoring it 4.33 out of 5. Government procurement influence scored a more measured 3.63. The authors read that gap as a signal: procurement linkage needs to be deliberately constructed rather than assumed to follow automatically from the leaderboard’s existence.
The paper acknowledges the survey’s limits honestly. The sample came from Nasscom’s professional network and skewed heavily toward production AI developers 52 of 82 respondents fell into that category. Civil society was nearly absent, with just four respondents identifying primarily with government or policy. End users of AI systems weren’t represented at all. The results offer useful evidence about practitioner preferences, not a representative picture of Indian society as a whole.
The Case Against the Critics
The paper directly addresses the most compelling objections to its position.
“Regional AI leaderboards will just be captured by local incumbents.” This is a real risk but it argues for careful governance design, not for abandoning governance altogether. Multi-stakeholder boards with no single dominant member, term limits for board positions, external audits, and published methodology all reduce capture risk significantly. An imperfect governance structure can be improved over time. The complete absence of governance cannot.
“Fragmentation will prevent meaningful comparison across regions.” The proposed answer is federation, not unification. A shared reporting schema allows results to be placed side by side without requiring identical scoring systems everywhere. A registry of regional AI leaderboards that meet baseline governance standards published conflict-of-interest policies, a documented appeals process, independent methodology review would provide the necessary foundation. An optional meta-view, hosted by a neutral body, could surface cross-regional comparisons without overriding regional rankings. MLCommons, the W3C, and ISO all operate on federated principles across independent national bodies. Cross-border standardization is achievable.
“This is just protectionism dressed up as equity advocacy.” Regional AI leaderboards don’t lock out global models from participating. They provide evaluation environments where performance for regional populations can be fairly measured. The paper argues explicitly that the goal is complementary infrastructure additional venues for honest assessment not replacement of existing systems.
“Governance failures affect Europe too why single out the Global South?” Yes, Basque and Welsh speakers face poorly calibrated AI systems too. But a Basque speaker has EU regulatory pressure, academic funding streams, and enterprise alternatives. A Santhali speaker or Hausa speaker has none of those fallback options. Commercial pressure creates accountability for the Global North. Governance is the only available accountability mechanism for everyone else.
Building It Right Before the Window Closes
The paper’s most urgent point may be about timing, not governance design.
Once an organization establishes itself as the de facto AI evaluator for a region, its rankings get written into procurement criteria, funding requirements, and research publication norms. Rebuilding governance into a captured institution is vastly harder than designing it correctly from the beginning. Regional AI leaderboards are being created right now, across India, Africa, and the Arab world. The window to build them with proper governance is open at this moment. The paper argues it won’t stay open long.
The minimum requirements for a trustworthy regional AI leaderboard are actually modest: an independent, multi-stakeholder governance board with term limits; a published conflict-of-interest policy with clear disclosure and recusal rules; a standardized submission protocol; a formal dispute resolution process; and a common reporting schema that allows results to be compared with other regional leaderboards.
For funding sustainability, the paper identifies four concrete paths: national AI-mission allocations (India’s $1.2 billion program is an obvious candidate); multi-donor consortia with governance firewalls between funders and operators; industry membership structures with disclosed contributions and limits on any single member’s voting power; and philanthropic anchor funding matched by public co-investment. NIST, MLCommons, and the W3C web standards body all demonstrate that multi-stakeholder evaluation infrastructure can sustain itself at scale.
For global leaderboard operators specifically, the asks are direct: publish formal conflict-of-interest policies, create formal appeals processes, expand to metrics suited to morphologically complex languages, add evaluation of code-switching performance, and include regional benchmarks from IndicSUPERB, IrokoBench, AlGhafa, and SEA-HELM in standard evaluation suites.
The Stakes Are Bigger Than Benchmarks
It would be easy to read a story about AI benchmark governance and conclude: this is a specialist technical debate for machine learning researchers. The stakes reach well past that.
India is spending $1.2 billion to bring AI into government services. A procurement officer selecting a speech recognition system for a citizen helpline will consult AI leaderboards. Those rankings will tell them which system performs best. If the rankings don’t reflect actual performance for Hindi, Tamil, or Telugu speakers or for someone speaking a regional variety of their language the deployed system will fail the very people it’s supposed to serve. Those users will have no clear explanation for why the government’s AI assistant doesn’t understand them. Research on algorithmic transference suggests they’ll also lose trust in the institution that deployed it.
The same pattern repeats in agriculture, where AI disease detection tools calibrated on North American crops get promoted to smallholder farmers across sub-Saharan Africa. In medicine, where cardiovascular risk models validated on Western populations get applied to South Asian patients without re-validation. In weather forecasting, where the regions most exposed to weather stress receive the least reliable AI predictions.
The paper’s central argument is that governance transforms a useful tool into a trusted institution. It creates accountability where accountability otherwise doesn’t exist. Commercial pressure does that job for the Global North. Without governance structures that require inclusion, the Global South receives the deployment of AI without the accountability and the failures, once documented, simply stay documented.
One line from the paper deserves to be remembered: “Imperfect representation under transparent governance is improvable. Exclusion under no governance is permanent.”
The rankings carry real consequences for real people. Getting them right for everyone, not just for those with commercial pull is not a technical problem waiting for a technical solution. It’s an institutional problem waiting for institutions willing to take it seriously.
The Above Article is based on the position paper “AI Leaderboards Are Underserving the Global South: A Case Study from India” by Sourav Banerjee (IIT Kharagpur and Shunya Labs) and Saikat Saha (Nasscom), presented at the 43rd International Conference on Machine Learning, Seoul, South Korea, 2026.



