Position: AI Leaderboards Are Underserving the Global South: A Case Study from India

arXiv cs.AI Papers

Summary

This position paper argues that AI leaderboards are structurally ill-suited to serving the Global South due to lack of independent governance and conflict-of-interest policies, using India as a case study to highlight institutional barriers and propose solutions.

arXiv:2608.18117v1 Announce Type: new Abstract: This position paper argues that AI leaderboards are structurally ill-suited to serving the Global South because they lack independent governance, conflict-of-interest policies, and mechanisms for metric evolution. The barrier is not missing data; high-quality regional benchmarks already exist: IndicSUPERB, MILU, and LAHAJA for India; IrokoBench for Africa; AlGhafa for Arabic. The barrier is institutional design. Global leaderboards do not include these benchmarks, and no governance mechanism compels them to do so. Commercial pressure corrects leaderboard failures when paying customers in the Global North are affected. The Global South lacks equivalent leverage. Without governance, failures affecting Hindi, Swahili, or Arabic speakers persist indefinitely as documented but unaddressed gaps. Using India as a case study (1.4 billion people, 22 scheduled languages, high-quality benchmarks, but no trusted aggregation), we report findings from a consultation with 58 AI practitioners showing consistent preference for formal governance and disclosure-based conflict management. The solution is not more data but better institutions: regional leaderboards with independent governance from the start.
Original Article
View Cached Full Text

Cached at: 08/20/26, 09:57 AM

# Position: AI Leaderboards Are Underserving the Global South: A Case Study from India
Source: [https://arxiv.org/html/2608.18117](https://arxiv.org/html/2608.18117)
###### Abstract

This position paper argues that AI leaderboards are structurally ill\-suited to serving the Global South because they lack independent governance, conflict\-of\-interest policies, and mechanisms for metric evolution\. The barrier is not missing data\. High\-quality regional benchmarks already exist: IndicSUPERB, MILU, and LAHAJA for India; IrokoBench for Africa; AlGhafa for Arabic\. The barrier is institutional design\. Global leaderboards do not include these benchmarks, and no governance mechanism compels them to do so\. Commercial pressure corrects leaderboard failures when paying customers in the Global North are affected\. The Global South lacks equivalent leverage\. Without governance, failures affecting Hindi, Swahili, or Arabic speakers persist indefinitely as documented but unaddressed gaps\. Using India as a case study \(1\.4 billion people, 22 scheduled languages, high\-quality benchmarks but no trusted aggregation\), we report findings from a consultation with 82 AI practitioners showing strong preferences for formal, non\-government governance and for disclosure\-based conflict management\. Our contribution is institutional, not technical: we arguethatandwhygovernance is the load\-bearing solution, with India as the worked case\.

AI Evaluation, Leaderboards, Multilingual AI, Global South, India, Governance, Cultural AI, Accent Diversity

## 1Introduction

AI leaderboards are influential signals in the AI ecosystem\. Governments reference them in procurement guidance\(Office of Management and Budget,[2024](https://arxiv.org/html/2608.18117#bib.bib29); Ravindran,[2025](https://arxiv.org/html/2608.18117#bib.bib30)\)\. Enterprises consult them for vendor selection\. Investors cite them in due diligence\. Research documents that companies invest heavily in benchmark performance, with estimates of hundreds of thousands of dollars in compute for top scores\(Eriksson et al\.,[2024](https://arxiv.org/html/2608.18117#bib.bib9)\)\. When a leaderboard declares a model “best,” that signal enters procurement shortlists and investment memos\. Because benchmarks define the implicit loss functions that guide model optimization, leaderboard governance shapes Machine Learning \(ML\) research and is a technical concern for the ML community\.

We argue that AI leaderboards, lacking conflict\-of\-interest \(COI\) policies and independent oversight, are structurally ill\-suited to serving the Global South, and that regional leaderboard infrastructure with independent governance is necessary for emerging AI ecosystems to mature\.

Terminology\.Abenchmarkis a dataset paired with an evaluation protocol and metrics; MMLU is an example\. Aleaderboardis a platform that aggregates benchmark results into rankings; the Open LLM Leaderboard is an example\. Bygovernancewe mean the institutional structures that determine how decisions are made, who makes them, and how conflicts are managed\. Aconflict of interest\(COI\) exists when the same parties who benefit from high rankings also control the evaluation process, for example when a model developer also designs the benchmark, adjudicates edge cases, or decides when metrics are updated\. This is not about malice\. It is about structural incentives that can bias outcomes even with good intentions\. Our critique targets leaderboard governance, not benchmark quality\. Excellent benchmarks exist, but they lack the institutional infrastructure that would ensure their results are aggregated and ranked by parties without conflicts of interest\.

Benchmark quality is not the bottleneck\. High\-quality regional benchmarks already exist: IndicSUPERB, MILU, and LAHAJA for India; IrokoBench for Africa\(Adelani et al\.,[2025](https://arxiv.org/html/2608.18117#bib.bib1)\); AlGhafa for Arabic\. Global leaderboards simply do not use them\. What is missing isinstitutional infrastructure: trusted aggregation with transparent governance\. The gap is institutional, not technical\.

The stakes are immediate\. India’s $1\.2 billion IndiaAI Mission\(Ministry of Finance,[2025](https://arxiv.org/html/2608.18117#bib.bib23)\)references global benchmarks that inadequately capture Indian linguistic reality\. Similar patterns exist across the Global South: from Nigeria to Indonesia, governments and enterprises rely on leaderboard signals that do not reflect their populations’ needs\. A consultation with 82 AI practitioners \(Section[6\.2](https://arxiv.org/html/2608.18117#S6.SS2)\) confirms demand for formal governance: every respondent endorsed some form of formal governance, 64% preferred non\-government stewardship, and 68% favored disclosure\-based conflict management over pre\-emptive exclusion\.

Our position is falsifiable: Section[9](https://arxiv.org/html/2608.18117#S9)specifies conditions under which regional infrastructure would be unnecessary\. Throughout, we critique governance structures rather than specific organizations\. The concentration of expertise across roles \(model development, benchmark creation, leaderboard operation\) is natural in nascent ecosystems\. The question is whether governance structures exist to manage this concentration transparently\.

Scope and limitations\.As a position paper, this work arguesthatregional leaderboard infrastructure with independent governance is necessary, nothowto implement it\. We do not provide implementation blueprints, funding models, or technical specifications; these require context\-specific deliberation beyond any single paper’s scope\. Our consultation \(n=82\) is illustrative, not representative; the sample skews toward production AI developers \(52 of 82\) and cannot speak for end users or civil society, the very groups we argue are underrepresented\. Analogies we employ \(the 10/90 gap, examination boards as discussed later\) illuminate structural dynamics but have limits\. We identify coordination challenges \(federation across regional leaderboards\) without solving them\. These are directions for future work, not gaps that invalidate the core position\.

Conflict of Interest Disclosure\.S\. Banerjee is a researcher at the Indian Institute of Technology Kharagpur and is also affiliated with Shunya Labs, an Indian voice\-AI company\. S\. Saha is affiliated with Nasscom, the Indian industry association whose AI stakeholder networks distributed the practitioner consultation, part of which is reported in Section[6\.2](https://arxiv.org/html/2608.18117#S6.SS2)\. No Shunya Labs model is evaluated in this paper, and neither author serves on the governance board of any leaderboard discussed here\. We disclose these affiliations in keeping with the governance norms this paper advocates\.

## 2Why Leaderboards Matter

AI leaderboards have evolved from academic scoreboards into critical infrastructure that shapes the AI ecosystem\. Understanding who depends on them reveals the stakes of their limitations\.

Leaderboards as implicit loss functions\.From an optimization perspective, leaderboards function as implicit loss functions for the AI development ecosystem\. The metrics a leaderboard privileges become the objectives that organizations optimize against\. When a leaderboard rewards Word Error Rate on English text, the optimization loop creates incentives that favor models excelling at English word boundaries but encoding representations poorly suited to agglutinative morphology\. The deployment failures we document are consistent with systematic optimization pressures, not mere mismatches \(Figure[1](https://arxiv.org/html/2608.18117#S2.F1)\)\.

Global Leaderboards\(static metrics: WER, MMLU\)Technical Gapsmetrics, accents, cultureGovernance GapsCOI, access, accountabilityAdd Context Metrics\(still static, no evolution\)Governancestill missingRegional InfrastructureContext\-AwareMetricsIndependentGovernanceAdaptiveCapacity\(requires deliberate design\)Figure 1:Evolution of the argument\. Global leaderboards use static metrics with governance gaps\. Adding context metrics helps but remains static and unaccountable\. Regional infrastructureenablesbut does not guarantee solutions: it must be deliberately designed with context\-aware metrics, independent governance, and adaptive capacity\.Table 1:Stakeholder Impact of Unreliable LeaderboardsTable[1](https://arxiv.org/html/2608.18117#S2.T1)summarizes the stakeholder ecosystem\.Governmentsreference global rankings in AI procurement and policy documents\(Ravindran,[2025](https://arxiv.org/html/2608.18117#bib.bib30), commentary\)\.Enterprisescannot rely on English\-centric leaderboards to predict code\-switched performance, the dominant pattern in multilingual contact centers globally\(Bhogale et al\.,[2023](https://arxiv.org/html/2608.18117#bib.bib5)\)\.Startupsbuilding for local markets have no prominent venue to demonstrate regional language capabilities\.

Critically, ordinary citizens bear the risks but have no voice\.A farmer in rural India interacting with a government voice assistant, or a health worker in Nigeria using an AI diagnostic tool, consumes AI systems selected through leaderboard\-influenced procurement, yet has no seat at governance tables\. Research on “algorithmic transference”\(Longoni et al\.,[2022](https://arxiv.org/html/2608.18117#bib.bib20)\)shows that faulty AI erodes trust in deploying institutions, not just technology\. Technical and governance failures reinforce each other: inappropriate metrics persist because those who set them face no consequences from those who suffer the failures\.

The consequences are documented\. GPT\-4 generates significantly more hallucinations in Hindi than English\(Das et al\.,[2025](https://arxiv.org/html/2608.18117#bib.bib7)\); 88% of AI\-generated stories for Indian contexts contain cultural inaccuracies\(Bhagat et al\.,[2025](https://arxiv.org/html/2608.18117#bib.bib4)\); leading ASR models perform significantly worse on minority dialects\(Harris et al\.,[2024](https://arxiv.org/html/2608.18117#bib.bib11)\); agentic AI performance drops from 60% under benchmark conditions to 25% in production\(Mehta,[2025](https://arxiv.org/html/2608.18117#bib.bib22)\)\. These are systematic outcomes when models optimized for global benchmarks encounter regional reality\.

## 3The Asymmetric Impact on Global South

Governance failures in AI leaderboards harm all regions but disproportionately impact the Global South\. This section explains why\.

### 3\.1The Exclusion Problem

Global leaderboards do not merely have governance problems\. They have content problems\. The benchmarks they use systematically exclude the Global South\.

Consider the HuggingFace Open ASR Leaderboard\(Srivastav et al\.,[2025](https://arxiv.org/html/2608.18117#bib.bib39)\)222[https://huggingface\.co/spaces/hf\-audio/open\_asr\_leaderboard](https://huggingface.co/spaces/hf-audio/open_asr_leaderboard), accessed 24 May 2026\.and its “Multilingual ASR Evaluation” section\. It includes five languages: German, French, Italian, Spanish, and Portuguese\. All five represent European language families, and the evaluation sets reflect European varieties rather than Brazilian Portuguese, West African French, or Latin American Spanish spoken by the majority of these languages’ users\. No Hindi \(600M\+ speakers\), no Arabic \(370M\+ speakers\), no Swahili \(100M\+ speakers\), no Indonesian \(200M\+ speakers\), no Bengali \(270M\+ speakers\)\. Regional benchmarks exist for these languages: IndicSUPERB and LAHAJA for Indian languages, IrokoBench for African languages, AlGhafa for Arabic\. The leaderboard simply does not use them\.

This pattern extends beyond speech recognition\. Research on MMLU found that 84\.9% of geography questions focus exclusively on North American or European regions\(Singh et al\.,[2025b](https://arxiv.org/html/2608.18117#bib.bib38)\)\. English comprises 43\.8% of Common Crawl training data despite being spoken by less than 20% of the world’s population\. Arabic, the fifth most spoken language globally, accounts for less than 1% of training data\. Over 2,000 African languages are largely neglected in AI models\(Nature Editorial,[2025](https://arxiv.org/html/2608.18117#bib.bib27)\)\.

Data scarcity and governance failure are complementary, not opposed\.Upstream training\-data imbalance \(such as Arabic at less than 1% of Common Crawl\) is a real problem that requires investment in data collection and curation\. The English share of the indexed web is itself roughly 50%, so a 43\.8% share in Common Crawl reflects an internet\-wide skew, not a benchmark design flaw\. Our argument is about what happens*after*such data exists\. High\-quality regional benchmarks \(IndicSUPERB, MILU, LAHAJA, IrokoBench, AlGhafa\) have been built but are still ignored by global leaderboards because no governance mechanism compels their inclusion\. The data gap and the governance gap require different interventions, and both are necessary\. Our paper addresses the second, not because the first is unimportant, but because it is already widely recognized while the governance failure is not\.

### 3\.2Why Markets Will Not Fix This

Global health R&D governance produces systematic underinvestment in diseases affecting the Global South, not through malice but through structures optimizing for different priorities\. The “10/90 gap” describes how only 10% of health research addresses conditions causing 90% of the global disease burden\(Røttingen et al\.,[2013](https://arxiv.org/html/2608.18117#bib.bib33)\)\. Cancer research attracts funding because patients in wealthy countries can pay for treatments\. Malaria research is underfunded because affected populations cannot\.

AI evaluation infrastructure exhibits analogous neglect\. When an English ASR model fails, enterprises escalate, contracts are at stake, and engineering resources are allocated to fix it\. When the same model fails on Hindi or Hausa, the failure is documented in release notes as a scope limitation, acknowledged but not prioritized\. The asymmetry is structural, not intentional\.

This analogy to neglected tropical diseases is instructive only and has obvious limits: AI models can be adapted through fine\-tuning at lower cost than new drug trials\(Xia et al\.,[2024](https://arxiv.org/html/2608.18117#bib.bib44)\)\. But the analogy holds where it matters: fine\-tuningcapabilitydoes not create fine\-tuningincentives\. Without market pressure or governance mandates, technical possibility does not translate into actual adaptation\.

Market pressure creates accountability for the Global North\. It does not create accountability for the Global South\. This is not a criticism of any organization\. It is a description of structural incentives\. No market force will push global leaderboards to adopt regional benchmarks for Santhali, Hausa, or Quechua speakers\. Without governance mandating inclusion, inclusion will not happen\.

### 3\.3Four Mechanisms of Asymmetric Impact

Escape route asymmetry\.Sophisticated actors in resource\-rich contexts can conduct independent evaluation or maintain internal benchmarking infrastructure\. Resource\-constrained actors must accept global signals as authoritative\. A procurement officer in Bangalore must rely on leaderboards that a counterpart in Seattle can verify independently through enterprise\-grade evaluation teams\.

Bug\-fix priority asymmetry\.When GPT\-4 hallucinates in English, enterprise customers file tickets and the issue receives engineering attention\. When it hallucinates at higher rates in Hindi\(Das et al\.,[2025](https://arxiv.org/html/2608.18117#bib.bib7)\), the behavior is characterized as expected performance variance for lower\-resource languages\. No institution currently has the mandate to treat Hindi performance degradation as a first\-class defect\.

Institutional redundancy\.The Global North has multiple institutions checking AI capability claims: academic peer review, enterprise procurement teams, regulatory bodies, consumer advocacy organizations\. The Global South often has single points of failure where a flawed leaderboard becomes the sole arbiter of quality\.

Optimization lock\-in\.Leaderboards function as implicit loss functions, and architectural decisions favoring English compound over time as deployment generates training data\. When leaderboards do not measure what matters to the Global South, models do not optimize for it\. The gap widens\.

### 3\.4Governance as the Mechanism for Inclusion

Where market forces are absent, governance structures become the primary pathway to inclusion\. For the Global South, transparent leaderboard governance serves the function that commercial pressure serves for the Global North: it creates accountability\. Without governance structures that mandate inclusion of regional benchmarks, regional languages, and regional contexts, inclusion will not happen\. The data exists\. Global leaderboards ignore it\. Markets will not change this\. Only governance can\.

## 4Systematic Failures of Global Leaderboards

Global AI leaderboards fail diverse regions through technical limitations compounded by governance failures that prevent correction\. The technical problems are solvable\. The governance problems are why they remain unsolved\.

### 4\.1Technical Failures

Metric mismatch\.For ASR, Word Error Rate \(WER\) was designed for English and tends to disadvantage morphologically rich languages\(Javed et al\.,[2023a](https://arxiv.org/html/2608.18117#bib.bib15)\)\. For LLMs, MMLU\(Hendrycks et al\.,[2021](https://arxiv.org/html/2608.18117#bib.bib12)\)and similar benchmarks embed Western assumptions\. Models ranking highly on MMLU show substantial degradation on India\-specific questions \(MILU,Verma et al\.[2025](https://arxiv.org/html/2608.18117#bib.bib42)\), African knowledge \(IrokoBench,Adelani et al\.[2025](https://arxiv.org/html/2608.18117#bib.bib1)\), and Arabic contexts \(AlGhafa,Almazrouei et al\.[2023](https://arxiv.org/html/2608.18117#bib.bib3)\)\. The IndicParam benchmark\(Maheshwari et al\.,[2025](https://arxiv.org/html/2608.18117#bib.bib21)\)quantifies this gap: as reported in that benchmark, the best\-performing model evaluated \(Gemini\-2\.5\) reaches only 58% accuracy on low\-resource Indic languages, with GPT\-4 at 45%, substantially below these models’ English performance\.A note on model vintage:the reported numbers reflect models available at the time of the cited evaluations\. The absence of frontier\-model results on regional benchmarks is itself the governance gap we identify\. No institution is mandated to keep these results current, and so even when newer models exist their behavior on Hindi, Yoruba, or Arabic remains unmeasured in any trusted public venue\.

Code\-switching and accent blindness\.Code\-mixing prevalence in multilingual societies reached 60% by 2020\(Sengupta et al\.,[2024](https://arxiv.org/html/2608.18117#bib.bib35)\), yet global benchmarks treat it as exceptional\. The LAHAJA benchmark\(Javed et al\.,[2024](https://arxiv.org/html/2608.18117#bib.bib17)\)reveals 15\-30% ASR performance degradation across Hindi regional accents\. For LLMs, studies find significantly higher hallucination rates in Hindi and Farsi compared to English, even for questions rooted in Indic contexts\.

Knowledge conflicts\.Global models trained on Western corpora exhibitparametric dogmatism\(Su et al\.,[2024](https://arxiv.org/html/2608.18117#bib.bib40)\): they prioritize training\-time knowledge over retrieval\-augmented local context, producing “cultural hallucinations\.” A model may override Indian constitutional provisions with Western legal precedents, or contradict local medical practices even when explicitly provided as context\. This is not a data quantity problem\. It is a fundamental tension between model priors and local ground truth\.

The pattern extends across modalities\.We focus on language\-model leaderboards because the governance gap is most documented there, but the same structural pattern holds for vision, medical, agricultural, and weather AI\. Chest X\-ray vision\-language models underdiagnose marginalized groups, with the highest error rates for Black female patients\(Yang et al\.,[2025](https://arxiv.org/html/2608.18117#bib.bib45)\)\. Western cardiovascular risk models misclassified 80% of 4,975 Indian first\-time heart\-attack patients as low/moderate risk\(Gupta et al\.,[2026](https://arxiv.org/html/2608.18117#bib.bib10)\)\. PlantVillage crop\-disease accuracy fell from 99\.35% on its North American benchmark\(Mohanty et al\.,[2016](https://arxiv.org/html/2608.18117#bib.bib24)\)to 49% on Tanzanian cassava in deployment\(Mrisho et al\.,[2020](https://arxiv.org/html/2608.18117#bib.bib26)\)\. AI weather models inherit ERA5 biases against tropical regions where Africa has one\-eighth the WMO\-recommended station density\(World Meteorological Organization,[2024](https://arxiv.org/html/2608.18117#bib.bib43); Mozaffari et al\.,[2026](https://arxiv.org/html/2608.18117#bib.bib25)\)\. In every case, no governance mechanism required validation on affected populations before deployment\. Appendix[D](https://arxiv.org/html/2608.18117#A4)gives the full evidence\.

Every technical problem above has a known fix\. Better metrics exist\. Regional benchmarks exist\. The question is why global leaderboards have not adopted them\. The answer is governance\.

### 4\.2Governance Failures

Structural COI in practice\.As defined in Section[1](https://arxiv.org/html/2608.18117#S1), a conflict of interest arises when the same parties who benefit from high rankings also control evaluation\. Resource scarcity naturally concentrates expertise: the same people who have the skills to build models also have the skills to build benchmarks and run leaderboards\. This concentration is expected in nascent ecosystems\. The question is whether governance structures exist to manage it\.

The structural pattern\.Several prominent leaderboards allow the same players to submit models and evaluate competitors\. The HuggingFace Open ASR Leaderboard illustrates this pattern \(Figure[2](https://arxiv.org/html/2608.18117#S4.F2)\): some co\-authors of the leaderboard itself\(Srivastav et al\.,[2025](https://arxiv.org/html/2608.18117#bib.bib39)\)also develop top\-ranked models on that leaderboard\(Sekoyan et al\.,[2025](https://arxiv.org/html/2608.18117#bib.bib34)\), related training datasets by overlapping author teams\(Koluguri et al\.,[2025](https://arxiv.org/html/2608.18117#bib.bib18)\), and core architectural components\(Rekesh et al\.,[2023](https://arxiv.org/html/2608.18117#bib.bib31)\), with no published COI policy or grievance redressal mechanism\.

Major Compute Provider EmployeesSubmit Models\(top\-ranked\)Evaluate Models\(own & others\)OpenASR LeaderboardNo COI PolicyNo Grievance RedressalFigure 2:Structural pattern illustrated by the HuggingFace Open ASR Leaderboard: employees of a major compute provider who co\-authored the leaderboard also develop top\-ranked models, with no published COI policy or grievance redressal mechanism\.Why transparency of methods is insufficient\.One might argue that if evaluation code is public, conflicts do not matter\. But as in academic peer review, where criteria are public yet recusal is still required, the concern is not falsification but influence overwhichmetrics are chosen,howedge cases are adjudicated, andwhenbenchmarks are updated\. Transparency ofexecutiondoes not address conflicts indesign\.

Selective access\.Singh et al\. \([2025a](https://arxiv.org/html/2608.18117#bib.bib37)\)documented that one provider evaluated 27 model variants privately before public release, and the top two providers received 39\.6% of arena evaluation data while 83 open\-weight models combined received only 29\.7%\. Unlike academic peer review, leaderboards have no formal disclosure requirements, no recusal processes, and no dispute resolution\.

Declining transparency\.The Foundation Model Transparency Index found that average transparency scores fell from 58 in 2024 to 40 in 2025\(Bommasani et al\.,[2025](https://arxiv.org/html/2608.18117#bib.bib6)\)\. Companies remain most opaque about training data, compute, and post\-deployment usage\. Transparency is getting worse, not better\.

Goodhart dynamics\.Goodhart’s Law states that “when a measure becomes a target, it ceases to be a good measure\.” AI leaderboards exhibit this dynamic at scale: models optimize for benchmark metrics, and the metrics progressively lose validity as measures of true capability\(Thomas & Uminsky,[2022](https://arxiv.org/html/2608.18117#bib.bib41)\)\. But Goodhart dynamics do not affect all populations equally\. A metric designed for English word boundaries will eventually lose validity\. But English speakers experience genuine capability improvementsbeforethe metric breaks\. Populations whose needs were never encoded in the metric experience only the costs of optimization \(gaming, contamination, resource allocation to benchmark performance\) without ever experiencing the benefits\.

## 5What Institutional Infrastructure Means

The problem is not benchmarks\. India has IndicSUPERB, MILU, LAHAJA\. Africa has IrokoBench\. The Arabic world has AlGhafa\. Southeast Asia has SEA\-HELM\(Nguyen et al\.,[2024](https://arxiv.org/html/2608.18117#bib.bib28)\)\. The data exists\. The problem is that benchmarks without institutions are just datasets\.

Consider what global leaderboards currently provide: a website, a ranking, and implicit trust that the operators are acting in good faith\. What they lack is any structure that wouldjustifythat trust\.Institutional infrastructurefills this gap:

- •Trusted aggregation:A neutral body that ranks models without conflicts of interest\. Without this, rankings reflect the interests of whoever controls the leaderboard\.
- •Governance:Published COI policies, disclosure requirements, recusal processes\. Without this, the same organization can create benchmarks, submit models, and declare winners\.
- •Accountability:Formal dispute resolution when rankings are contested\. Without this, developers who believe they were unfairly evaluated have no recourse\.
- •Evolution:Mechanisms to adopt new metrics as understanding advances\. Without this, benchmarks ossify even as the field learns their limitations\.

High\-stakes examinations demonstrate that such structures are achievable at scale\. India’s JEE \(1\.5 million aspirants annually\) and China’s Gaokao maintain pre\-exam security while providing post\-hoc transparency: answer keys, scoring rubrics, and formal dispute resolution\. This infrastructure is imperfect, but failures trigger investigations and reforms\. AI leaderboards currently have none of this\.

Why must these institutions be built from inception? Path dependency\. Once an organization establishes itself as the de facto evaluator, its rankings become entrenched in procurement criteria, investment decisions, and research benchmarks\. Retrofitting governance onto captured infrastructure is far harder than designing it correctly from the start\. The Global South has a narrow window: regional leaderboards are emerging now\. The choice is whether they emerge with governance or without it\.

Why regional scale matters for governance\.The argument is not that regional institutions are immune to capture, but that regional scale changes accountability dynamics\. First,proximity: regional stakeholders can hold regional institutions accountable more directly \(a Hindi speaker has more leverage over an Indian institution than a global one\)\. Second,visibility: with fewer players, capture is harder to hide; everyone knows who the major actors are\. Third,correction: reforming a regional institution is more feasible than reforming an entrenched global one\. Fourth,aligned incentives: regional institutions derive legitimacy from serving regional populations, creating structural pressure toward inclusion that global institutions, responsive primarily to commercial pressure from the Global North, lack\.

## 6Case Study: India

India provides an ideal case study because it exhibits the core dynamics at maximum scale: 1\.4 billion people, extreme linguistic diversity \(22 scheduled languages, 80\+ Hindi dialects\), mature regional benchmarks developed by world\-class research groups, and significant government investment \($1\.2B IndiaAI Mission\)\. Crucially, India has the technical infrastructure \(high\-quality benchmarks exist\) but lacks the institutional infrastructure for trusted aggregation\. This is precisely the gap our position identifies\. We use India for depth; Section[6\.4](https://arxiv.org/html/2608.18117#S6.SS4)validates that the same pattern appears in Africa and the Arabic world\.

### 6\.1Existing Indic Benchmarks

Indian researchers have developed high\-quality benchmarks addressing gaps in global evaluation:

Table 2:Existing Indic AI Benchmarks \(non\-exhaustive\)The technical infrastructure exists\. What is missing istrusted leaderboard aggregationwith appropriate governance\.

### 6\.2Stakeholder Consultation Findings

Nasscom distributed a structured online survey on Indic AI leaderboard design through its stakeholder networks between December 2025 and March 2026, collecting 82 responses from a practitioner\-heavy pool\.333Conducted by Nasscom for this position work\. Full instrument, demographics, and graphical findings with statistical tests are in Appendices[A](https://arxiv.org/html/2608.18117#A1)and[B](https://arxiv.org/html/2608.18117#A2)\.The community judged the regional\-leaderboard proposal viable \(mean 4\.06/5\)\. Two findings stand out\. First, every respondent endorsed some form of formal governance; no one chose to leave governance ad hoc\. Second, within formal options, 64% preferred non\-government stewardship \(32\.9% Nasscom\-led, 31\.7% independent non\-profit\), versus 21% government\-driven and 9\.8% academia\-only \(χ2=26\.2\\chi^\{2\}=26\.2, df=4,p<10−4p<10^\{\-4\}vs\. uniform null\)\. On conflict management, 68% favored disclosure and recusal over pre\-emptive exclusion \(95% Wilson CI \[57\.6%, 77\.4%\],p≈1\.2×10−3p\\approx 1\.2\\times 10^\{\-3\}\)\.

Important limitations:convenience sample from Nasscom’s networks, skewed toward production AI developers \(52 of 82, 63%\) and toward respondents already engaged with governance questions\. Civil\-society and end\-user voices are underrepresented; only 4 of 82 respondents identified primarily with government/policy, so public\-sector views are likewise thin\. We treat this as illustrative triangulation, not representative evidence\. The full demographic breakdown is in Appendix[A](https://arxiv.org/html/2608.18117#A1)\.

Table 3:Stakeholder Consultation Key Findings \(n=82, Dec 2025 – Mar 2026\)Table[3](https://arxiv.org/html/2608.18117#S6.T3)summarizes key findings\. Beyond the headline numbers, thepatternof specific preferences is the actionable contribution: a clear tilt toward non\-government stewardship paired with disclosure\-based conflict management, a community willing to trade some reproducibility for dynamic anti\-gaming, and a quarterly refresh cadence\. The community’s strongest endorsement is for the leaderboard becoming a credible standard \(rated 4\.33 out of 5\); the most cautious dimension is procurement influence \(3\.63 out of 5\), which we treat as evidence that procurement linkage must be deliberately designed rather than assumed\.

### 6\.3Emerging Regional Infrastructure

Several regional leaderboard initiatives have emerged, validating demand for such infrastructure while illustrating the governance complexity we argue must be addressed early\.

The concentration pattern\.AI4Bharat \(IIT Madras\) launched the Indic LLM Arena in November 2025\(AI4Bharat,[2025](https://arxiv.org/html/2608.18117#bib.bib2)\)\. Figure[3](https://arxiv.org/html/2608.18117#S6.F3)illustrates the pattern: AI4Bharat creates benchmarks \(IndicSUPERB, MILU, LAHAJA\), builds models \(IndicWhisper, Airavata\), and operates the leaderboard, with funding from government, philanthropy, and industry\.

AI4Bharat\(IIT Madras\)BenchmarksModelsLeaderboardGovernmentPhilanthropyIndustryCommercialVentures\(co\-founders\)Figure 3:AI4Bharat illustrates how regional ecosystems naturally develop: a single organization creates benchmarks, builds models, and operates the leaderboard, with diverse funding sources\. This is not criticism but illustration of why governance frameworks become necessary as such entities scale\.Talent concentration is afeatureof nascent ecosystems, not a flaw unique to any organization\. This is precisely why formal COI policies become prerequisites for trust as entities scale\. Research shows thatpublicdisclosure eliminates negative impact on trust\(Liu et al\.,[2023](https://arxiv.org/html/2608.18117#bib.bib19)\)\.

Nasscom\-led32\.9%Independent non\-profit31\.7%Government\-driven20\.7%Academia consortium9\.8%Other4\.9%Figure 4:Governance model preference \(n=82\)\. All five options are forms of formal governance; no respondent chose “no governance”\. Combined non\-government stewardship \(Nasscom \+ Independent non\-profit\) accounts for 64% of preferences\.
### 6\.4Pattern Validation: Africa and Arabic

The India case is not unique\.Africa:IrokoBench\(Adelani et al\.,[2025](https://arxiv.org/html/2608.18117#bib.bib1)\)provides high\-quality evaluation for 16 African languages, yet global leaderboards do not use it\. Over 2,000 African languages remain largely neglected\(Nature Editorial,[2025](https://arxiv.org/html/2608.18117#bib.bib27)\)\. The governance gap is identical: benchmarks exist, trusted aggregation does not\.Arabic world:AlGhafa\(Almazrouei et al\.,[2023](https://arxiv.org/html/2608.18117#bib.bib3)\)and the Open Arabic LLM Leaderboard\(El Filali et al\.,[2025](https://arxiv.org/html/2608.18117#bib.bib8)\)address Arabic evaluation, but face the same concentrated expertise pattern where benchmark creators also submit models\. Arabic, the fifth most spoken language globally, accounts for less than 1% of training data\. In both cases, the technical infrastructure exists; institutional infrastructure does not\. India is the detailed case; Africa and Arabic validate the pattern\.

## 7Alternative Views

We address seven counterarguments to our position\.

### 7\.1“Regional Leaderboards Risk Capture and Gaming”

Argument: Regional leaderboards may replicate governance failures at smaller scale, with local incumbents capturing governance\.

Response: These concerns argue for careful governance design, not against governance\. For capture risk: multi\-stakeholder boards, term limits, and external audits\. For gaming: quarterly refresh with contamination detection\.

### 7\.2“Fragmentation Will Prevent Interoperability”

Argument: Regional leaderboards with different metrics will prevent meaningful cross\-regional comparison\.

Response: Federation, not unification, is the answer\. Cross\-regional comparison does not require identical leaderboards; it requires standardized result schemas so that per\-language, per\-dialect performance can be reported in a common schema and inspected side\-by\-side, while acknowledging that scores across different benchmarks are not apples\-to\-apples even when the metric \(e\.g\., WER\) is the same\. A practical federation has three layers: \(1\) a shared schema for reporting per\-language, per\-dialect, per\-domain performance \(analogous to MLPerf’s submission format\); \(2\) a registry of trusted regional leaderboards that conform to baseline governance standards \(COI policy, appeals, public methodology\); \(3\) an optional meta\-view that surfaces cross\-regional comparisons without overriding regional rankings\. The MLCommons benchmarking consortium, the W3C web standards process, and the ISO certification ecosystem are existence proofs that federated coordination across independent bodies is feasible\. We acknowledge that the protocol specification itself is future work\(Rodriguez Müller & Schade,[2024](https://arxiv.org/html/2608.18117#bib.bib32)\), and Appendix[E](https://arxiv.org/html/2608.18117#A5)sketches a starting architecture\. Transparent diversity serves users better than false uniformity, particularly when uniformity is achieved by excluding the contexts that matter most to the Global South\.

### 7\.3“Governance Is a Global Problem”

Argument: European languages also suffer from governance failures\. Isn’t this universal?

Response: Yes, but impact is asymmetric\. A Basque speaker has ecosystem redundancy: EU regulatory pressure, academic funding, enterprise alternatives\. A Santhali or Hausa speaker has no alternative pathway\. Market pressure creates accountability for the Global North; governance is the Global South’s only mechanism for inclusion\.

### 7\.4“Organizations Participate in Good Faith, and Regional Advocacy Is Just Nationalism”

Argument: Organizations contribute positively to leaderboards; questioning their involvement is unfair, and regional\-leaderboard advocacy is disguised protectionism\.

Response: We question governance structures, not intentions\. Resource scarcity naturally concentrates expertise across roles \(build models, create benchmarks, run leaderboards\)\. This is expected in nascent ecosystems\. The question is whether governance exists to manage this concentration\. Individual good faith does not substitute for institutional accountability, and network effects create winner\-take\-all dynamics where building alternative credibility takes years\(Singh et al\.,[2025a](https://arxiv.org/html/2608.18117#bib.bib37)\)\. On the nationalism objection: regional leaderboards do not prevent global models from participating\. They provide evaluation contexts where regional performance can be fairly assessed\. The goal is complementary infrastructure, not replacement\.

### 7\.5“Funding Risks and Governance Details Unspecified”

Argument: Regional leaderboards require sustained funding, and governance processes remain unspecified\.

Response: Evaluation infrastructure is a public good\. The question “who pays?” is downstream of governance: it cannot be answered without first deciding who bears costs, on whose behalf, and accountable to whom\. Saying cost matters more than governance is like saying a budget matters more than a constitution\. NIST \(publicly funded\), MLCommons \(multi\-stakeholder consortium with industry dues and academic in\-kind\), and W3C \(member dues plus host institutions\) are existence proofs\. Four sustainability paths are concrete: \(a\) national AI\-mission allocations \(India’s $1\.2B IndiaAI Mission; similar Indonesian, Brazilian, and African Union initiatives\); \(b\) multi\-donor consortia with governance firewalls; \(c\) industry membership tiers with disclosed contributions and capped voting power; \(d\) philanthropic anchor funding matched against public co\-investment\. Procurement linkage strengthens sustainability: ISO certification raises developing\-country exports by 44\.9% on average\(International Organization for Standardization,[2024](https://arxiv.org/html/2608.18117#bib.bib14)\)\. Our respondents are cautious about procurement uptake \(3\.63 out of 5\), which we treat as evidence that procurement linkage requires deliberate design, not a reason to defer governance\. Appendix[C](https://arxiv.org/html/2608.18117#A3)details the governance framework\.

### 7\.6“End\-User Voice Remains Absent”

Argument: The paper invokes affected citizens but proposes governance that underweights their representation\.

Response: This critique is valid, and argues for stronger governance, not against it\. Governance frameworks create thepossibilityof inclusion through user panels and complaint processes\. Imperfect representation under transparent governance is improvable; exclusion under no governance is permanent\.

## 8Call to Action

Our position implies specific actions\. A minimum viable regional leaderboard requires: \(1\) independent governance with a multi\-stakeholder board and term limits; \(2\) a published COI policy; \(3\) a standardized submission protocol; \(4\) dispute resolution; \(5\) interoperability via common result schemas\. Sustainable funding models include public funding, industry consortia with governance firewalls, or hybrid approaches\. Governance framework details are provided in Appendix[C](https://arxiv.org/html/2608.18117#A3)\.

For nations and funding bodies:Invest in regional leaderboard infrastructure as a public good with multi\-year funding commitments\. Aggregate existing benchmarks under transparent governance\. Require independence criteria for funded evaluation efforts\.

For global leaderboard maintainers:Adopt formal COI policies with disclosure and recusal\. Publish testing access policies\. Create formal appeals processes\. Expand to morphologically appropriate metrics and code\-switching evaluation\. Include regional benchmarks from the Global South\.

For researchers and industry:Recognize regional benchmarks as authoritative for their domains, not secondary to global leaderboards\. Require multilingual evaluation for claims of general capability\. Support multiple evaluation pathways and contribute to regional benchmark development\.

## 9Conclusion

AI leaderboards are infrastructure\. They determine which systems are considered state\-of\-the\-art and enter procurement and investment decisions\. When this infrastructure fails to serve a region, the AI ecosystem in that region operates on unreliable signals\.

We have argued that AI leaderboards, absent COI policies, independent oversight, and mechanisms for metric evolution, are structurally ill\-suited to serving the Global South\. The problem is not data\. Regional benchmarks exist\. The problem is institutional: no governance structure adopts, maintains, and evolves these improvements\. Market pressure corrects failures in the Global North; the Global South has no equivalent leverage\. Governance is the substitute\.

Our position is falsifiable\.If global leaderboards demonstrate reliable evaluation with governance that emerging participants trust, regional infrastructure would be unnecessary\. The operationalizable conditions are: \(1\)Technical:report language\-family\-appropriate metrics, include code\-switching evaluation for top\-10 multilingual markets, disaggregate by regional accent, and incorporate culturally\-grounded knowledge benchmarks; \(2\)Inclusive:include regional benchmarks \(IndicSUPERB, IrokoBench, AlGhafa, SEA\-HELM\) in standard evaluation suites; \(3\)Adaptive:maintain formal processes for adopting new metrics as research evolves; \(4\)Governance:publish explicit COI policies, with more than 50% of surveyed Global South practitioners rating governance as “trustworthy”; \(5\)Access:keep data\-access disparity between top\-3 commercial providers and median open\-weight models at most 2x; \(6\)Impact parity:governance failures affect Global North and Global South practitioners comparably\.

Until these conditions are met, regional leaderboards are not fragmentation of the evaluation landscape but necessary infrastructure\. The choice is not between global coordination and regional autonomy but between functional infrastructure and continued underservice of the Global South\.

## Acknowledgements

We thank the ICML 2026 anonymous reviewers and our area chair for substantive engagement during the review and discussion period that materially strengthened this paper, including the suggestions on cross\-regional coordination, on funding feasibility, and on the relationship between training\-data scarcity and evaluation governance\.

## Disclaimer

The views expressed in this paper are those of the authors in their personal capacity and do not represent the official positions of IIT Kharagpur, Shunya Labs, or Nasscom\. This is an argumentative position paper, not an empirical study; the practitioner consultation reported here is illustrative rather than representative, and its limitations are documented in Section[6\.2](https://arxiv.org/html/2608.18117#S6.SS2)and Appendix[A](https://arxiv.org/html/2608.18117#A1)\.

## References

- Adelani et al\. \(2025\)Adelani, D\. I\. et al\.IrokoBench: A new benchmark for african languages in the age of large language models\.In*Proceedings of NAACL*, 2025\.URL[https://aclanthology\.org/2025\.naacl\-long\.139/](https://aclanthology.org/2025.naacl-long.139/)\.
- AI4Bharat \(2025\)AI4Bharat\.Indic LLM\-Arena: A crowd\-sourced evaluation platform for Indian languages\.AI4Bharat, IIT Madras, November 2025\.URL[https://ai4bharat\.iitm\.ac\.in/blog/indic\-llm\-arena](https://ai4bharat.iitm.ac.in/blog/indic-llm-arena)\.
- Almazrouei et al\. \(2023\)Almazrouei, E\., Cojocaru, R\., Baldo, M\., Malartic, Q\., Alobeidli, H\., Mazzotta, D\., Penedo, G\., Campesan, G\., Farooq, M\., Alhammadi, M\., Launay, J\., and Noune, B\.AlGhafa evaluation benchmark for Arabic language models\.In*Proceedings of ArabicNLP*, 2023\.URL[https://aclanthology\.org/2023\.arabicnlp\-1\.21/](https://aclanthology.org/2023.arabicnlp-1.21/)\.
- Bhagat et al\. \(2025\)Bhagat, K\., Bhatt, S\., Velagapudi, A\., Vashistha, A\., Dave, S\., and Pruthi, D\.TALES: A taxonomy and analysis of cultural representations in LLM\-generated stories, 2025\.URL[https://arxiv\.org/abs/2511\.21322](https://arxiv.org/abs/2511.21322)\.88% of AI\-generated stories contain cultural inaccuracies; errors more prevalent in low\-resource languages\.
- Bhogale et al\. \(2023\)Bhogale, K\., Javed, T\., Raman, A\., et al\.Vistaar: Diverse benchmarks and training sets for Indian language ASR\.In*Proceedings of Interspeech*, 2023\.URL[https://arxiv\.org/abs/2305\.15386](https://arxiv.org/abs/2305.15386)\.
- Bommasani et al\. \(2025\)Bommasani, R\., Klyman, K\., Kapoor, S\., Maslej, N\., Longpre, S\., Xiong, B\., Liang, P\., and Wan, A\.Foundation model transparency index 2025\.Technical report, Stanford University, 2025\.URL[https://crfm\.stanford\.edu/fmti/paper\.pdf](https://crfm.stanford.edu/fmti/paper.pdf)\.Comprehensive evaluation of foundation model transparency practices\.
- Das et al\. \(2025\)Das, A\. et al\.Investigating hallucination in conversations for low resource languages, 2025\.URL[https://arxiv\.org/abs/2507\.22720](https://arxiv.org/abs/2507.22720)\.GPT\-4o generates significantly more hallucinations in Hindi than other languages\.
- El Filali et al\. \(2025\)El Filali, A\., Aloui, M\., Husaain, T\., Alzubaidi, A\., Boussaha, B\. E\. A\., Cojocaru, R\., Fourrier, C\., Habib, N\., and Hacid, H\.The Open Arabic LLM Leaderboard 2\.Hugging Face Blog, February 2025\.URL[https://huggingface\.co/blog/leaderboard\-arabic\-v2](https://huggingface.co/blog/leaderboard-arabic-v2)\.
- Eriksson et al\. \(2024\)Eriksson, M\., Purificato, E\., et al\.Can we trust AI benchmarks? an interdisciplinary review of current issues in AI evaluation\.In*Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society*, 2024\.URL[https://ojs\.aaai\.org/index\.php/AIES/article/view/36595](https://ojs.aaai.org/index.php/AIES/article/view/36595)\.Documents companies spending heavily on benchmark scores; OpenAI estimated hundreds of thousands on compute for benchmarks\.
- Gupta et al\. \(2026\)Gupta, R\., Mohan, I\., Narula, J\., and Sharma, K\. K\.Recalibration of cardiovascular risk prediction models for the Indian population: A cohort study of first\-time acute myocardial infarction patients\.*Indian Heart Journal*, 2026\.Cohort of 4,975 first\-time AMI patients in India; Framingham, ACC/AHA ASCVD, and WHO charts misclassified nearly 80% as low or moderate risk\.
- Harris et al\. \(2024\)Harris, C\. et al\.Minority English dialects vulnerable to automatic speech recognition inaccuracy\.Georgia Institute of Technology, November 2024\.URL[https://www\.gatech\.edu/news/2024/11/15/minority\-english\-dialects\-vulnerable\-automatic\-speech\-recognition\-inaccuracy](https://www.gatech.edu/news/2024/11/15/minority-english-dialects-vulnerable-automatic-speech-recognition-inaccuracy)\.SAE significantly outperforms AAVE, Spanglish, and Chicano English on leading ASR models\.
- Hendrycks et al\. \(2021\)Hendrycks, D\., Burns, C\., Basart, S\., Zou, A\., Mazeika, M\., Song, D\., and Steinhardt, J\.Measuring massive multitask language understanding\.In*International Conference on Learning Representations*, 2021\.URL[https://openreview\.net/forum?id=d7KBjmI3GmQ](https://openreview.net/forum?id=d7KBjmI3GmQ)\.
- IISc Machine Intelligence and Language Engineering Lab\(2024\) \(MILE\)IISc Machine Intelligence and Language Engineering \(MILE\) Lab\.IISc\-MILE ASR corpora for Indian languages\.Indian Institute of Science, Bangalore, 2024\.URL[https://mile\.ee\.iisc\.ac\.in/](https://mile.ee.iisc.ac.in/)\.Lab releases include Tamil and Kannada ASR corpora used in IndicSUPERB and related benchmarks\.
- International Organization for Standardization \(2024\)International Organization for Standardization\.Economic benefits of standards: ISO methodology 2\.0 and country case studies\.Technical report, ISO, 2024\.URL[https://www\.iso\.org/publication/PUB100440\.html](https://www.iso.org/publication/PUB100440.html)\.Synthesis across country studies finds standards\-related exports rising by an average of 44\.9% in developing economies adopting international conformity assessment\.
- Javed et al\. \(2023a\)Javed, T\., Bhogale, K\., Raman, A\., Kunchukuttan, A\., Kumar, P\., and Khapra, M\. M\.IndicSUPERB: A speech processing universal performance benchmark for Indian languages\.In*Proceedings of AAAI*, 2023a\.URL[https://arxiv\.org/abs/2208\.11761](https://arxiv.org/abs/2208.11761)\.
- Javed et al\. \(2023b\)Javed, T\., Joshi, S\., Bhogale, K\., et al\.Svarah: Evaluating English ASR systems on Indian accents\.In*Proceedings of Interspeech*, 2023b\.URL[https://arxiv\.org/abs/2305\.15760](https://arxiv.org/abs/2305.15760)\.
- Javed et al\. \(2024\)Javed, T\., Nawale, J\., Joshi, S\., George, E\., Bhogale, K\., Mehendale, D\., and Khapra, M\. M\.LAHAJA: A robust multi\-accent benchmark for evaluating Hindi ASR systems\.In*Proceedings of Interspeech*, 2024\.URL[https://arxiv\.org/abs/2408\.11440](https://arxiv.org/abs/2408.11440)\.
- Koluguri et al\. \(2025\)Koluguri, N\. R\., Sekoyan, M\., Zelenfroynd, G\., Meister, S\., Ding, S\., Kostandian, S\., Huang, H\., Karpov, N\., Balam, J\., Lavrukhin, V\., Peng, Y\., Papi, S\., Gaido, M\., Brutti, A\., and Ginsburg, B\.Granary: Speech recognition and translation dataset in 25 European languages, 2025\.URL[https://arxiv\.org/abs/2505\.13404](https://arxiv.org/abs/2505.13404)\.Training dataset for top\-ranked models; first author also co\-authors leaderboard\.
- Liu et al\. \(2023\)Liu, F\., AlShebli, B\., and Rahwan, T\.Current policies governing editorial conflicts of interest are ineffective\.*arXiv preprint arXiv:2307\.00794*, 2023\.URL[https://arxiv\.org/abs/2307\.00794](https://arxiv.org/abs/2307.00794)\.Analysis of 500K papers: public COI disclosure eliminates negative impact on reader trust\.
- Longoni et al\. \(2022\)Longoni, C\., Cian, L\., and Kyung, E\. J\.Algorithmic transference: People overgeneralize failures of AI in the government\.*Journal of Marketing Research*, 60\(1\):170–188, 2022\.doi:10\.1177/00222437221110139\.URL[https://journals\.sagepub\.com/doi/full/10\.1177/00222437221110139](https://journals.sagepub.com/doi/full/10.1177/00222437221110139)\.Empirical study: AI failures are generalized more broadly than human failures due to out\-group homogeneity perceptions\.
- Maheshwari et al\. \(2025\)Maheshwari, A\., Sharma, K\., Patel, V\., and Maheshwari, A\.IndicParam: A human\-curated benchmark for evaluating LLMs on low\-resource Indic languages, 2025\.URL[https://arxiv\.org/abs/2512\.00333](https://arxiv.org/abs/2512.00333)\.Top model \(Gemini\-2\.5\) achieves only 58% accuracy on low\-resource Indic languages; GPT\-4 at 45%\.
- Mehta \(2025\)Mehta, S\.CLEAR: A multi\-dimensional framework for evaluating enterprise agentic AI systems, 2025\.URL[https://arxiv\.org/abs/2511\.14136](https://arxiv.org/abs/2511.14136)\.Agent performance drops from 60% to 25% in consistency tests; 50x cost variation for similar accuracy\.
- Ministry of Finance \(2025\)Ministry of Finance\.Union Budget 2025\-26: Expenditure budget volume ii\.Technical report, Government of India, February 2025\.URL[https://www\.indiabudget\.gov\.in/doc/eb/sbe27\.pdf](https://www.indiabudget.gov.in/doc/eb/sbe27.pdf)\.
- Mohanty et al\. \(2016\)Mohanty, S\. P\., Hughes, D\. P\., and Salathé, M\.Using deep learning for image\-based plant disease detection\.*Frontiers in Plant Science*, 7:1419, 2016\.doi:10\.3389/fpls\.2016\.01419\.URL[https://www\.frontiersin\.org/articles/10\.3389/fpls\.2016\.01419/full](https://www.frontiersin.org/articles/10.3389/fpls.2016.01419/full)\.Reports 99\.35% accuracy on PlantVillage lab benchmark for crop disease detection\.
- Mozaffari et al\. \(2026\)Mozaffari, A\., Lim, W\., Patel, A\., and Karpatne, A\.Regional bias in AI weather forecasting: Tropical and Global South performance of ERA5\-trained models, 2026\.URL[https://arxiv\.org/abs/2602\.04571](https://arxiv.org/abs/2602.04571)\.AI weather models \(GraphCast, Pangu\-Weather, FourCastNet\) inherit ERA5 observational biases against tropical regions\.
- Mrisho et al\. \(2020\)Mrisho, L\. M\., Mbilinyi, N\. A\., Ndalahwa, M\., Ramcharan, A\. M\., Kehs, A\. K\., McCloskey, P\. C\., Murithi, H\., Hughes, D\. P\., and Legg, J\. P\.Accuracy of a smartphone\-based object detection model, PlantVillage nuru, in identifying the foliar symptoms of cassava diseases in Tanzania\.*Frontiers in Plant Science*, 11:590889, 2020\.doi:10\.3389/fpls\.2020\.590889\.URL[https://www\.frontiersin\.org/articles/10\.3389/fpls\.2020\.590889/full](https://www.frontiersin.org/articles/10.3389/fpls.2020.590889/full)\.Field deployment of PlantVillage model for cassava disease detection in Tanzania achieved 49% accuracy in farmer field conditions\.
- Nature Editorial \(2025\)Nature Editorial\.AI models are neglecting African languages: scientists want to change that\.*Nature*, July 2025\.doi:10\.1038/d41586\-025\-02292\-5\.URL[https://www\.nature\.com/articles/d41586\-025\-02292\-5](https://www.nature.com/articles/d41586-025-02292-5)\.Over 2,000 African languages remain unrepresented in major AI benchmarks\.
- Nguyen et al\. \(2024\)Nguyen, X\.\-P\., Aljunied, M\., Joty, S\., and Bing, L\.SEA\-HELM: A holistic evaluation of large language models for Southeast Asian languages\.Salesforce AI Research, 2024\.URL[https://github\.com/SeaLLMs/SEA\-HELM](https://github.com/SeaLLMs/SEA-HELM)\.Benchmark for evaluating LLMs on Southeast Asian languages\.
- Office of Management and Budget \(2024\)Office of Management and Budget\.Advancing the responsible acquisition of AI in government\.Technical Report M\-24\-18, Executive Office of the President, October 2024\.URL[https://www\.whitehouse\.gov/wp\-content/uploads/2024/10/M\-24\-18\-AI\-Acquisition\-Memorandum\.pdf](https://www.whitehouse.gov/wp-content/uploads/2024/10/M-24-18-AI-Acquisition-Memorandum.pdf)\.Federal AI procurement guidance requiring evaluation frameworks for vendor selection\.
- Ravindran \(2025\)Ravindran, B\.How India can build inclusive, culturally relevant language models\.*Nature India*, January 2025\.doi:10\.1038/d44151\-025\-00084\-4\.URL[https://www\.nature\.com/articles/d44151\-025\-00084\-4](https://www.nature.com/articles/d44151-025-00084-4)\.
- Rekesh et al\. \(2023\)Rekesh, D\., Koluguri, N\. R\., Kriman, S\., Majumdar, S\., Noroozi, V\., Huang, H\., Hrinchuk, O\., Puvvada, K\., Kumar, A\., Balam, J\., et al\.FastConformer with linearly scalable attention for efficient speech recognition\.In*IEEE Automatic Speech Recognition and Understanding Workshop \(ASRU\)*\. IEEE, 2023\.URL[https://arxiv\.org/abs/2305\.05084](https://arxiv.org/abs/2305.05084)\.Core architecture of top\-ranked models; co\-authors overlap with leaderboard maintainers\.
- Rodriguez Müller & Schade \(2024\)Rodriguez Müller, P\. and Schade, S\.Interoperability assessments: Exploring expected benefits, efforts and challenges\.Technical report, European Commission, 2024\.URL[https://publications\.jrc\.ec\.europa\.eu/repository/handle/JRC137063](https://publications.jrc.ec.europa.eu/repository/handle/JRC137063)\.Cross\-border interoperability requires sustained institutional investment, continuous capability building, and cultural/organizational changes\.
- Røttingen et al\. \(2013\)Røttingen, J\.\-A\., Regmi, S\., Eide, M\., Young, A\. J\., Viergever, R\. F\., Årdal, C\., Guzman, J\., Edwards, D\., Matlin, S\. A\., and Terry, R\. F\.Mapping of available health research and development data: what’s there, what’s missing, and what role is there for a global observatory?*The Lancet*, 382\(9900\):1286–1307, 2013\.doi:10\.1016/S0140\-6736\(13\)61046\-6\.Established the “10/90 gap”: less than 10% of global health research addresses problems of 90% of world’s population\.
- Sekoyan et al\. \(2025\)Sekoyan, M\., Koluguri, N\. R\., Tadevosyan, N\., Żelasko, P\., Bartley, T\., Karpov, N\., Balam, J\., and Ginsburg, B\.Canary\-1b\-v2 & parakeet\-tdt\-0\.6b\-v3: Efficient and high\-performance models for multilingual ASR and AST, 2025\.URL[https://arxiv\.org/abs/2509\.14128](https://arxiv.org/abs/2509.14128)\.Top\-ranked ASR models on Open ASR Leaderboard; authors overlap with leaderboard maintainers\.
- Sengupta et al\. \(2024\)Sengupta, A\., Das, S\., Akhtar, M\. S\., and Chakraborty, T\.Social, economic, and demographic factors drive the emergence of Hinglish code\-mixing on social media\.*Humanities and Social Sciences Communications*, 11:606, 2024\.doi:10\.1038/s41599\-024\-03058\-6\.URL[https://www\.nature\.com/articles/s41599\-024\-03058\-6](https://www.nature.com/articles/s41599-024-03058-6)\.
- Singh et al\. \(2024\)Singh, H\., Gupta, N\., Bharadwaj, S\., Tewari, D\., and Talukdar, P\.IndicGenBench: A multilingual benchmark to evaluate generation capabilities of LLMs on Indic languages\.In*Proceedings of ACL*, 2024\.URL[https://aclanthology\.org/2024\.acl\-long\.595/](https://aclanthology.org/2024.acl-long.595/)\.
- Singh et al\. \(2025a\)Singh, S\., Nan, Y\., Wang, A\., D’souza, D\., Kapoor, S\., Üstün, A\., Koyejo, S\., Deng, Y\., Longpre, S\., Smith, N\. A\., Ermis, B\., Fadaee, M\., and Hooker, S\.The leaderboard illusion\.In*Advances in Neural Information Processing Systems: Datasets and Benchmarks Track*, 2025a\.URL[https://openreview\.net/forum?id=4Ae8edNqm0](https://openreview.net/forum?id=4Ae8edNqm0)\.
- Singh et al\. \(2025b\)Singh, S\. et al\.Global MMLU: Understanding and addressing cultural and linguistic biases in multilingual evaluation, 2025b\.URL[https://arxiv\.org/abs/2412\.03304](https://arxiv.org/abs/2412.03304)\.Only 15% of MMLU has been culturally adapted for non\-Western contexts\.
- Srivastav et al\. \(2025\)Srivastav, V\., Zheng, S\., Bezzam, E\., Le Bihan, E\., Moumen, A\., and Gandhi, S\.Open ASR leaderboard: Towards reproducible and transparent multilingual and long\-form speech recognition evaluation, 2025\.URL[https://arxiv\.org/abs/2510\.06961](https://arxiv.org/abs/2510.06961)\.Leaderboard co\-authored by researchers from major compute providers who also develop top\-ranked models\.
- Su et al\. \(2024\)Su, Z\., Zhang, J\., Qu, X\., Zhu, T\., Li, Y\., Sun, J\., Li, J\., Zhang, M\., and Cheng, Y\.ConflictBank: A benchmark for evaluating knowledge conflicts in large language models\.In*Advances in Neural Information Processing Systems*, 2024\.URL[https://arxiv\.org/abs/2408\.12076](https://arxiv.org/abs/2408.12076)\.
- Thomas & Uminsky \(2022\)Thomas, R\. L\. and Uminsky, D\.Reliance on metrics is a fundamental challenge for AI\.*Patterns*, 3\(5\):100476, 2022\.doi:10\.1016/j\.patter\.2022\.100476\.URL[https://www\.cell\.com/patterns/fulltext/S2666\-3899\(22\)00056\-4](https://www.cell.com/patterns/fulltext/S2666-3899(22)00056-4)\.“Current AI approaches have weaponized Goodhart’s law by centering on optimizing a particular measure as a target\.”\.
- Verma et al\. \(2025\)Verma, S\., Khan, M\. S\. U\. R\., Kumar, V\., Murthy, R\., and Sen, J\.MILU: A multi\-task Indic language understanding benchmark\.In*Proceedings of NAACL*, 2025\.URL[https://aclanthology\.org/2025\.naacl\-long\.507/](https://aclanthology.org/2025.naacl-long.507/)\.
- World Meteorological Organization \(2024\)World Meteorological Organization\.State of the global climate 2024: Observing\-network gaps in Africa and small island developing states\.Technical Report WMO\-No\. 1347, WMO, 2024\.URL[https://library\.wmo\.int/idurl/4/68585](https://library.wmo.int/idurl/4/68585)\.Africa operates at roughly one\-eighth of the WMO\-recommended surface\-station density\.
- Xia et al\. \(2024\)Xia, Y\., Kim, J\., Chen, Y\., Ye, H\., Kundu, S\., Hao, C\., and Talati, N\.Understanding the performance and estimating the cost of LLM fine\-tuning, 2024\.URL[https://arxiv\.org/abs/2408\.04693](https://arxiv.org/abs/2408.04693)\.Fine\-tuning enables task specialization using limited compute resources in a cost\-effective manner compared to pretraining\.
- Yang et al\. \(2025\)Yang, Y\., Zhang, H\., Gichoya, J\. W\., Katabi, D\., and Ghassemi, M\.Demographic bias of vision\-language foundation models in medical imaging\.*Science Advances*, 11\(4\), 2025\.URL[https://www\.science\.org/doi/10\.1126/sciadv\.adq0305](https://www.science.org/doi/10.1126/sciadv.adq0305)\.Vision\-language models consistently underdiagnose marginalized groups in chest X\-ray analysis, with highest error rates for Black female patients\.

## Appendix

## Appendix ASurvey Instrument and Methodology

### A\.1Purpose

The consultation was conducted to gather expert design input for an Indic AI leaderboard\. The survey was distributed by Nasscom \(the Indian National Association of Software and Service Companies\) through its AI stakeholder networks between December 2025 and March 2026\. The instrument was developed iteratively with input from industry policy experts and structured around four areas: governance and architecture; technical challenges and solutions; implementation and operations; roadmap and success criteria\.

### A\.2Survey Instrument

All questions permitted multiple selections unless noted as single\-select\.

Q1\. Language selection \(multi\)\.Starting set for the Indic AI Leaderboard\.

Q2\. Governance model \(single\)\.Nasscom hosts/governs \(hybrid funding, industry\-academia advisory\); Academia consortium hosts/governs \(government funding\); Independent non\-profit hosts/governs \(hybrid funding\); Government\-driven \(MeitY or IndiaAI aegis, PPP funding\); Other\.

Q3\. Conflict management \(single\)\.Pre\-emptive exclusion vs\. disclosure/recusal\.

Q4\. ASR datasets \(multi\)\.Vistaar, LAHAJA, IISc\-MILE, Svarah, ULCA/Bhashini\.

Q5\. ASR metrics \(multi\)\.WER, CER, KER, plus open write\-in\.

Q6\. LLM benchmarks \(multi\)\.MILU, IndicGenBench, IndicMMLU, IndicGLUE, M3LS, IndicXTREME, DRISHTIKON, MultiQ, ParamBench\.

Q7\-Q9\. LLM metric families \(multi\)\.Semantic similarity \(BERTScore, BLEURT, COMET, BARTScore, PRISM, MoverScore\); factuality \(FactCC/DAE/QAGS, SummaC\); bias/cultural \(Stereotype/Bias Score, Cultural Accuracy Score, Indic\-Bias, SANSKRITI\)\.

Q10\. Evaluation method \(single\)\.Human, LLM, or hybrid\.

Q11\. Anti\-gaming strategy \(single\)\.Static public, static hidden, dynamic, or adversarial\.

Q12\. Submission requirements \(multi\)\.API documentation, technical specifications, reliability, security\.

Q13\. Submission cadence \(single\)\.Quarterly or half\-yearly\.

Q14\. Emerging priorities \(multi\)\.VLMs, multimodal, code generation, agentic AI, domain\-specific\.

Q15\. Success criteria \(1\-5 scale, single\)\.Used at 2 years; influence on procurement; credibility; ASR/LLM track parity\.

Q16\-Q17\. Open\-ended\.Additional suggestions\.

Demographics \(multi, with consent\)\.Stakeholder category, expertise areas, affiliation\. No verification of respondent identity\.

### A\.3Sampling and Limitations

Convenience sample drawn from Nasscom’s networks\. This introduces selection bias toward practitioners with existing engagement in the association’s network\. Self\-reported expertise and affiliations\. The final survey window closed at 82 responses\. Civil\-society and end\-user voices remain underrepresented; only 4 of 82 respondents \(4\.9%\) identified primarily with the “Government/Policy” stakeholder category, which limits our ability to disaggregate public\-sector views even though the option was offered in the instrument\.

### A\.4Note on Survey Evolution \(n=58 to n=82\)

The conference submission cited an interim snapshot of n=58 collected through January 2026\. The survey window remained open and continued to collect responses through March 2026, yielding the final n=82 dataset reported in this camera\-ready version\. All percentages, charts, and headline findings here use the final n=82 data\. The change of sample size did not alter the direction of any reported finding\. The largest shift between snapshots was the disclosure\-versus\-exclusion preference \(71% at n=58, 68% at n=82\), which remains a clear majority\. We report this evolution transparently in keeping with the governance norms this paper advocates\.

## Appendix BSelected Practitioner\-Consultation Findings \(n=82\)

This appendix presents selected findings only\. A comprehensive version of the underlying consultation is scheduled for release by Nasscom AI in early July 2026\.

This appendix presents the selected Nasscom consultation findings as bar charts, each paired with a brief statistical test\. Unless stated otherwise, percentages are computed overn=82n=82, 95% confidence intervals \(CIs\) are Wilson intervals, andpp\-values come from exact binomial tests against the stated null or from Pearsonχ2\\chi^\{2\}goodness\-of\-fit tests against a uniform distribution over the offered options\. Likert items \(1–5 scale\) are tested with one\-samplett\-tests against a neutral mean of 3, using a conservative assumed standard deviation of 1\.0\.

Common plot conventions\.All bar charts use raw response counts on thexx\-axis \(oryy\-axis for horizontal bars\), withn=82n=82unless the item is multi\-select \(in which case totals exceed 82 and we label the chart accordingly\)\. Where a clear “null” alternative exists we annotate the chart with the relevant test statistic\.

### B\.1Respondent Demographics

The respondent pool was practitioner\-heavy\. Stakeholder category and area\-of\-expertise were both multi\-select questions, so totals exceed 82\.

02020404060605252252518181313665544Selections \(total=129\)Prod\. AI Dev\.Enterprise ConsumerIndep\. ConsultantAcademic/ResearchInvestor/VCCivil Soc\./NGOGovernment/Policy\(a\)Stakeholder categories \(total==129 selections\)\.020204040606062623636343434342222161699Selections \(total=219\)Product Dev\.ML/AI Infra\.LLMsAI Ethics/Fair\.Policy/Gov\.Indian Lang\.ASR Systems\(b\)Expertise areas \(total==219 selections\)\.
Figure 5:Respondent demographics\. \(a\) Production AI developers \(52\) are roughly twice the next\-largest group, so reported preferences should be read primarily as the views of builders and enterprise deployers; Government/Policy \(4\) and Civil Society/NGO \(5\) are the smallest structured groups, and 6 free\-text responses \(e\.g\., startup founders, EdTech\) bring the total to 129 selections\. \(b\) ML/AI infrastructure, LLMs, and AI ethics form a tight second tier behind product development, indicating a technically grounded but practitioner\-skewed pool; Indian\-language linguistics \(16\) and ASR systems \(9\) are the smallest structured tiers, and 6 free\-text write\-ins bring the total to 219\.Limitations\.The convenience sample over\-represents builders \(Production AI Developer \+ Enterprise AI Consumer \+ Independent Consultant==95 of 129 selections, 74%\) and under\-represents civil society \(5/129, 3\.9%\) and government/policy \(4/129, 3\.1%\)\. Findings should be read as views of an engaged practitioner community, not a population\-representative survey\.

### B\.2Language Coverage \(Q1\)

055101015152020252530303535404045455050555560606565707075758080SanskritUrduPunjabiOdiaAssameseMalayalamGujaratiKannadaBengaliMarathiTeluguTamilHindiInd\. English1212101010109988171719192020222227273434353565657171Respondents endorsing language \(n=82n=82, multi\-select\)\(a\)Top tier \(counts≥8\\geq 8\)\.00\.50\.5111\.51\.5222\.52\.5333\.53\.5444\.54\.5555\.55\.5666\.56\.5777\.57\.588BodoDogriNepaliMaithiliManipuriSantaliSindhiKashmiriKonkani222222333333334455Respondents endorsing language \(n=82n=82, multi\-select\)\(b\)Long tail \(counts≤5\\leq 5\)\.
Figure 6:Language coverage priorities\. Indian English \(71\) and Hindi \(65\) are clear anchors\. Six languages clear the 22\-respondent \(\>\>25%\) threshold: Indian English, Hindi, Tamil, Telugu, Marathi, Bengali\. The long tail \(17 of 23 listed languages with<<25% endorsement\) signals demand for phased but broad eventual coverage\. Sanskrit \(12\) over\-indexes relative to its spoken\-population share, reflecting the academic sub\-pool\.
### B\.3Governance Preference \(Q2\)

Nasscom\-ledIndep\. NPGovernmentAcademiaOther0101020203030uniform null==16\.42727262617178844Respondents \(n=82n=82\)Figure 7:Governance model preference\. Non\-government stewardship \(Nasscom\-led\+\+independent non\-profit\)==53/82, 64\.6% \(95% CI \[53\.8%, 74\.1%\]\)\. Distribution is significantly non\-uniform:χ2=26\.2\\chi^\{2\}=26\.2, df=4=4,p=2\.9×10−5p=2\.9\\times 10^\{\-5\}\. All four structured options are formal\-governance models; no respondent chose “no governance”\. The two community\-anchored options jointly carry the modal weight\.
### B\.4Conflict Management \(Q3\)

Disclosure/RecusalPre\-emptive Exclusion020204040606056562626Respondents \(n=82n=82\)Figure 8:Conflict\-of\-interest management preference\. Disclosure/recusal==56/82, 68\.3% \(95% CI \[57\.6%, 77\.4%\]\)\. Two\-sided binomial test vs\. a 50/50 null:p=1\.2×10−3p=1\.2\\times 10^\{\-3\}\. The community prefers managed participation over exclusion\.
### B\.5ASR Datasets and Metrics \(Q4, Q5\)

01010202030304040505060607070IISc\-MILELAHAJASvarahULCA/BhashiniVistaar33334545474751516363Respondents endorsing dataset \(n=82n=82, multi\-select\)\(a\)Datasets: Vistaar \(63/82, 77%\) anchors; four clear 50%\.WERCERKER02020404060608080696947474040Respondents \(n=82n=82\)\(b\)Metrics: WER \(84%\), CER \(57%\), KER \(49%\)\.
Figure 9:ASR datasets and metrics\. \(a\) Vistaar \(63/82, 77%\) is the natural anchor; four datasets clear 50% endorsement, indicating expectation of a multi\-dataset ASR suite rather than a single benchmark\. \(b\) WER is dominant; CER and KER are widely endorsed as complements rather than alternatives\. Ten open\-text write\-ins requested semantic\-similarity and token\-error variants, signalling that standard rate\-based metrics are necessary but not sufficient for Indic\-language nuance\.
### B\.6LLM Datasets \(Q6\) and Metric Families \(Q7–Q9\)

0551010151520202525303035354040454550505555606065657070IndicMMLUDRISHTIKONIndicGenBenchMILU4141414150505959Respondents endorsing dataset \(n=82n=82, multi\-select\)Figure 10:LLM datasets\. MILU \(59/82, 72%\) leads; no dataset clears 75% endorsement, confirming that no single dataset covers language breadth, cultural depth, domain knowledge, and generation capability simultaneously\. A composite suite is expected\.01010202030304040505060607070MoverScorePRISMBARTScoreBLEURTCOMETBERTScore272730303333343435356363Respondents \(n=82n=82, multi\-select\)Figure 11:Semantic\-similarity metrics for LLMs\. BERTScore \(77%\) dominates; five secondary metrics cluster tightly between 33% and 43%, indicating expectation of a multi\-metric semantic suite rather than a single canonical score\.Stereotype/BiasCultural Accuracy020204040606059595656Respondents \(n=82n=82\)Figure 12:Bias and cultural\-fairness metrics\. The two scores are statistically indistinguishable \(Stereotype/Bias 72%, Cultural Accuracy 68%; difference of proportionsp=0\.62p=0\.62\)\. Both are treated as essential, reflecting the community’s view that Indic AI evaluation must explicitly test for social fairness, not only task accuracy\. Factuality items \(FactCC/DAE/QAGS family, 67/82, 82%\) are the clear leader within that separate question\.
### B\.7Evaluation Method \(Q10\)

Hybrid LLM\+HumanLLM JudgesHuman Judges0202040406060uniform null==27\.36262121288Respondents \(n=82n=82\)Figure 13:Evaluation method preference\. Hybrid \(LLM\+human\)==62/82, 75\.6% \(95% CI \[65\.3%, 83\.6%\]\)\. Test against a uniform 3\-option null:χ2=66\.2\\chi^\{2\}=66\.2, df=2=2,p<10−14p<10^\{\-14\}\. Hybrid evaluation reflects awareness of both LLM\-judge biases and human\-judge scalability limits\.
### B\.8Anti\-Gaming Strategy \(Q11\)

DynamicStatic PublicAdversarialStatic Hidden020204040uniform null==19\.2546461616111144Respondents \(n=77n=77answered\)Figure 14:Anti\-gaming strategy \(5 of 82 respondents abstained\)\. Dynamic evaluation==46/82, 56\.1% \(95% CI \[45\.3%, 66\.3%\]\); among the 77 who answered, Dynamic was 59\.7%\. Test against a uniform 4\-option null over the 77 answered responses:χ2=53\.3\\chi^\{2\}=53\.3, df=3=3,p<10−10p<10^\{\-10\}\. Even at the cost of strict reproducibility, the community prefers leaderboards that can rotate test items to resist memorization and overfitting\.
### B\.9Submission Requirements and Cadence \(Q12, Q13\)

0101020203030404050506060707080809090ReliabilitySecurityTech\. SpecsAPI Docs5656636370707777Respondents \(n=82n=82, multi\-select\)Figure 15:Submission requirements\. All four standards\-style requirements clear 68% endorsement \(lowest is Reliability at 56/82, 68%; highest is API Documentation at 94%\), indicating broad alignment that submissions should meet production\-grade engineering standards before evaluation\.QuarterlyHalf\-yearly020204040606061612121Respondents \(n=82n=82\)Figure 16:Submission cadence\. Quarterly windows==61/82, 74\.4% \(95% CI \[64\.0%, 82\.6%\]\)\. Two\-sided binomial test vs\. a 50/50 null:p=1\.1×10−5p=1\.1\\times 10^\{\-5\}\. Three of four respondents want four submission windows per year, balancing freshness against operational burden\.
### B\.10Emerging Technology Priorities \(Q14\)

020204040606066666161606059593333Respondents \(n=82n=82\)Specialized Dom\.MultimodalAgentic AIVisual LanguageCode Gen\.

Figure 17:Emerging\-technology priorities\. Specialized domains \(healthcare, legal, agriculture, education\)==66/82, 80\.5% \(95% CI \[70\.6%, 87\.6%\]; vs\. 50% null:p=2\.3×10−8p=2\.3\\times 10^\{\-8\}\)\. Multimodal \(74%\) and agentic \(73%\) are statistically indistinguishable from each other\. Code generation \(40%\) trails noticeably, suggesting the community sees Indic\-language code as a secondary near\-term priority\.
### B\.11Success Criteria \(Q15, 1–5 Scale\)

334455neutral==34\.334\.334\.214\.214\.064\.063\.633\.63Mean rating \(1–5\)Credible StandardUsed in 2 YearsASR/LLM ParityProcurement Inf\.

Figure 18:Success criteria, mean Likert rating \(n=82n=82, neutral==3\)\. All four items are significantly above neutral \(one\-samplett\-tests with assumed SD==1\): Credible Standardt=12\.1t=12\.1,p<10−15p<10^\{\-15\}; Used in 2 Yearst=11\.0t=11\.0,p<10−15p<10^\{\-15\}; ASR/LLM Parityt=9\.6t=9\.6,p<10−15p<10^\{\-15\}; Procurement Influencet=5\.7t=5\.7,p<10−7p<10^\{\-7\}\. The community is most confident in trust building \(4\.33\) and most cautious on government procurement uptake \(3\.63\), suggesting that policy linkage is a follower variable rather than a precondition\.
### B\.12What the Statistical Inference Adds

The bar charts above show that the headline percentages reported in Section[6\.2](https://arxiv.org/html/2608.18117#S6.SS2)are not artifacts of small samples or balanced splits\. For every multi\-option question, the response distribution is significantly non\-uniform atp<10−4p<10^\{\-4\}or better, with the modal choice having a 95% CI that does not include the chance level\. For every dichotomous question, the majority preference is significant atp<10−3p<10^\{\-3\}\. For every Likert success criterion, the mean is significantly above neutral atp<10−7p<10^\{\-7\}\. The consultation cannot establish population\-level claims \(the sample is a convenience pool, not a probability sample\), but within that pool the preferences are sharp, consistent, and unlikely under any plausible chance\-only model\.

## Appendix CIllustrative Governance Framework

This appendix presents one possible instantiation of governance structures, not a prescriptive specification\. Regional implementations should adapt these elements to local institutional contexts, legal frameworks, and stakeholder needs\.

Table 4:Board Composition and TermsTable 5:Conflict of Interest Policy
## Appendix DCross\-Modal Governance Gap

The body \(Section[4\.1](https://arxiv.org/html/2608.18117#S4.SS1)\) summarizes evidence that the governance gap extends beyond language models\. This appendix documents the cases in more detail\.

Medical diagnostics\.A 2025 study of vision\-language models in chest X\-ray diagnosis found that models consistently underdiagnose marginalized groups, with the highest error rates for Black female patients\(Yang et al\.,[2025](https://arxiv.org/html/2608.18117#bib.bib45)\)\. In cardiology, a study of 4,975 first\-time heart\-attack patients in India found that the widely used Western cardiovascular risk models \(Framingham, ACC/AHA ASCVD, WHO charts\) misclassified nearly 80% of patients as low or moderate risk, despite the patients having presented with acute infarction\(Gupta et al\.,[2026](https://arxiv.org/html/2608.18117#bib.bib10)\)\. The models were never re\-validated against South Asian physiology and lipid profiles before being adopted as clinical reference points\.

Agriculture\.The PlantVillage crop\-disease detection model reported 99\.35% accuracy on its North American lab benchmark\(Mohanty et al\.,[2016](https://arxiv.org/html/2608.18117#bib.bib24)\)\. When the same model class was deployed for cassava disease detection in Tanzania, accuracy fell to 49%\(Mrisho et al\.,[2020](https://arxiv.org/html/2608.18117#bib.bib26)\)\. The drop reflects field\-versus\-lab image conditions, regional disease variants, and crop varieties absent from training data\. No governance mechanism required validation on the deployment population before the model was promoted to smallholder farmers\.

Weather prediction\.Modern AI weather models such as Pangu\-Weather, GraphCast, and FourCastNet are trained on the ECMWF ERA5 reanalysis, which itself is constructed from a global observation network in which Africa has roughly one\-eighth of the WMO\-recommended surface\-station density\(World Meteorological Organization,[2024](https://arxiv.org/html/2608.18117#bib.bib43)\)\. AI weather models inherit this observational bias, producing higher forecast errors over tropical regions\(Mozaffari et al\.,[2026](https://arxiv.org/html/2608.18117#bib.bib25)\)\. Communities most exposed to climate volatility are the least well served by the systems being deployed to manage it\.

The common pattern\.In every case the technical fix is known \(collect more local data, run domain\-adapted fine\-tuning, re\-validate against local populations\)\. The bottleneck is institutional: no governance mechanism mandates validation on the affected populations before deployment, and no appeals process exists when deployed systems fail\. A better maker is not a substitute for a checker\.

## Appendix EFederation Architecture Sketch

Section[7\.2](https://arxiv.org/html/2608.18117#S7.SS2)argues that federation, not unification, is the answer to cross\-regional comparability\. This appendix sketches a starting architecture\.

Layer 1: Shared result schema\.A common JSON schema for reporting per\-language, per\-dialect, per\-domain, and per\-task performance, analogous to the MLPerf submission format\. The schema is descriptive \(“Model X scored 94% CER on Hindi LAHAJA dev split, run on date Y, hardware Z, prompt template P”\), not prescriptive about which metrics to optimize\. Regional leaderboards remain free to weight metrics differently in their headline rankings\.

Layer 2: Trusted\-leaderboard registry\.A lightweight registry of regional leaderboards that conform to baseline governance standards: a published COI policy, a documented appeals process, a public methodology, and an independent fairness\-audit panel\. Conformance is self\-declared with periodic peer review \(analogous to ICANN’s accreditation of registrars or ISO’s accreditation of national standards bodies\)\.

Layer 3: Optional meta\-view\.A meta\-leaderboard, hosted by a neutral body \(such as MLCommons or a UN\-affiliated entity\), surfaces cross\-regional comparisons in the shared schema without overriding regional rankings\. The meta\-view exposes capability claims to multilingual scrutiny without imposing a global ranking\.

Why this works\.The three layers separate three concerns that the current ecosystem conflates: \(1\) what to measure \(regional autonomy\), \(2\) how to report measurements \(shared schema\), and \(3\) who can claim trust \(registry\)\. Federation across MLCommons, NIST, W3C, and ISO operates on similar principles\. Protocol specification is future work and beyond the scope of this position\.

Similar Articles

On the missing benchmarks layer and a potential solution

arXiv cs.AI

This paper argues that Latin America lacks a benchmark layer for native AI development and proposes an open, task-first EvalsHub infrastructure, with LatamBoard as its first regional instance, to audit AI systems and direct optimization toward local needs.

High-stakes game of musical chairs!

Reddit r/singularity

The author analyzes the current AI race, arguing that big corporations are using high costs to outlast smaller competitors, but predicts a shift to flat-fee pricing and locally-run AI in the long term, advocating for government co-built data centers for sovereign AI infrastructure.