On the missing data layer and a potential solution

arXiv cs.AI Papers

Summary

This paper argues that Latin America lacks the datasets and benchmarks layers needed to build its own AI, and proposes DataHub, an open, incentive-driven platform with a task-first ontology to index and contribute regional datasets.

arXiv:2608.02949v1 Announce Type: new Abstract: Latin America is missing two foundational layers of AI infrastructure: the dataset layer and the benchmark layer. This paper targets the dataset layer. The dataset layer faces two compounding problems: discovery and supply. Latin American AI datasets exist but are scattered across platforms with no shared index. Even with perfect indexing, the total volume would remain far below what frontier AI development requires. We propose DataHub: a task-first data infrastructure organized through the ontology /<task?>/<domain?>/<language?>, with mechanisms for dataset discovery, metadata, contribution, licensing, and reuse.
Original Article
View Cached Full Text

Cached at: 08/05/26, 07:38 AM

# On the missing data layer and a potential solution
Source: [https://arxiv.org/html/2608.02949](https://arxiv.org/html/2608.02949)
\(July 29, 2026\)

###### Abstract

Latin America is missing two foundational layers required to build its own AI: adatasets layerand abenchmarks layer\. These are not optional infrastructure — they are the minimum prerequisites for both building and evaluating AI systems\. Their absence carries a dual cost\. On the industry side, most AI tasks are universal but the data needed to get state\-of\-the\-art performance in a given context is regionally peculiar\. On the public\-institution side, the cost is different in kind: models trained elsewhere encode foreign value systems, languages, and historical narratives\[tao2024cultural\], shaping how citizens are represented, served, and governed\. Together, the absence of both layers constitutes a loss in both industrial productivity and sovereign capacity over a technology that is becoming critical infrastructure\. The dataset layer faces two compounding problems: discovery and supply\. Latin American datasets exist but are scattered across places like Hugging Face and the appendices of research papers, with no shared index\. Even pooled, their volume falls far below what frontier AI development requires\. The two problems reinforce each other: with nowhere to publish for return, fewer datasets get released; with few releases, no critical mass forms\. The proposed DataHub is designed to break this loop from both ends — indexing what exists and making contribution a visible, rewarded act\. The DataHub is organized around a task\-first ontology with three nested levels:task\(e\.g\. transcription\), thendomain\(e\.g\. medical\), thenlanguage and regional variant—– \(e\.g\. Rioplatense Spanish\)\. Each level compounds the value of the last, paying off on both sides: for industry, in performance on the target task; for public institutions and the populations they serve, in representation and fit\. We prefer amultipolar AI worldand believe Latin America should be one of those poles, contributing on its own terms\. The DataHub must therefore be open by design — because collaboration matters most for regions not at the frontier — and incentive\-driven by construction, since a commons only stays alive when contributors are rewarded\. Open questions remain: regulatory heterogeneity across jurisdictions, shared standards for metadata and licensing, collaboration norms, and the institutional incentives needed to unlock data currently kept internal\. This paper is a kickstart by the SURUS team on the missing data layer; and an invitation to companies, institutions, and individuals across the region to shape what the foundational infrastructure for Latin American AI should become, and to contribute and benefit from it\. We prefer a multipolar Arica should be one ofthose poles, contributing on its own terms\. The Hub must therefore be open by design — because collabons not at the frontier —and incentive\-driven by construction, since a commons only stays alive when contributors are rewarded\. Open questions remain: regulatory heterogeneity across jurisdictions, shared standards for metadata and licensing, collaboration norms, and needed to unlock datacurrently kept internal\. This paper is a kickstart by the SURUS team — a working artifact targeting the dataset layer; the benchmark layer is the subject of LatamBoard, athey are an invitation toresearchers, institutionegion to shape what thefoundational infrastructure for Latin American AI should become\.

1SURUS2Independent

Contact:francis@surus\.lat

Keywords:AI, Data Infrastructure, Latin America, Dataset Discovery, Multipolar AI

###### Contents

1. [1Introduction](https://arxiv.org/html/2608.02949#S1)
2. [2Two missing layers, two lanes of cost](https://arxiv.org/html/2608.02949#S2)
3. [3State of Latin American datasets](https://arxiv.org/html/2608.02949#S3)
4. [4A task\-first ontology for AI](https://arxiv.org/html/2608.02949#S4)
5. [5Open by design, incentive\-driven by construction](https://arxiv.org/html/2608.02949#S5)
6. [6Open problems](https://arxiv.org/html/2608.02949#S6)
7. [7A kickstart and an invitation](https://arxiv.org/html/2608.02949#S7)
8. [References](https://arxiv.org/html/2608.02949#bib)

SURUS Team · Buenos Aires · 2026

## 1Introduction

The state of AI in Latin America today is one of consumption\. The region uses models built elsewhere, trained on data collected elsewhere, evaluated against tasks defined elsewhere\[[18](https://arxiv.org/html/2608.02949#bib.bib36)\]\. For genuinely universal and verifiable capabilities — coding, general\-purpose reasoning, mathematics — this is workable; outside those, performance degrades on tasks where regional context \(represented through data\) matters, and no evaluation relevant to the region exists to measure how much\.

Two foundational layers are missing: thedataset layer, which training requires, and thebenchmark layer, which evaluation requires\. Without datasets, the region cannot train its own models, agents, or downstream systems\. Without benchmarks, it cannot evaluate anything it builds — nor evaluate the foreign systems already in use\.

This paper targets the dataset layer; the benchmark layer is addressed in LatamBoard, a companion effort by the SURUS team\. It covers the cost of both missing layers, the current state of Latin American datasets, the Hub’s task\-first ontology, the proposed open design, the geopolitical position from which the design follows, and the open problems that remain\.

## 2Two missing layers, two lanes of cost

These two layers are not optional infrastructure — they are theminimum prerequisitesfor both building and evaluating AI systems\. Without datasets, the region cannot train models, agents, or downstream systems — any of them\. Without benchmarks, it cannot evaluate what it builds, and — just as critically — cannot evaluate the foreign systems already deployed across its institutions and industries\. Both absences are active today\.

Without a benchmark layer, industries cannot evaluate the models they procure — and adopted models can underperform silently on the very tasks they were bought to perform\. Regional peculiarities matter for most industrially useful tasks; without benchmarks on those tasks, procurement decisions are made blind\. The failure is silent: the model returns an output, but not the performance needed for full adoption and diffusion on the tasks that matter\[[12](https://arxiv.org/html/2608.02949#bib.bib18)\]\.

For public institutions, the absence of benchmarks forecloses something larger: the possibility of becomingindependent auditorsof both regional and worldwide AI systems in use\. Some universities independently measure poverty or inequality, producing formal, region\-anchored evaluations that their countries and the region use\. Latin America’s universities and public institutions in general could produce the same for AI: formal evaluations of what models — local and global — get right and wrong for the populations they serve\.

On the industry side, the problem is performance: the task is often universal, but the data needed to perform it at the state\-of\-the\-art isregionally peculiar\. A pest\-classification model trained on European or North American agriculture will not reliably identify\[[2](https://arxiv.org/html/2608.02949#bib.bib34)\]the pests that determine harvests in Argentina or Brazil\. The task is universal; the crops, the species, and the visual distribution are not\. This is the performance gap in concrete form\.

On the public\-institution side, the cost is different in kind: models trained elsewhere encode foreign value systems, languages, and historical narratives — shaping how citizens are represented, served, and governed\. The failure mode is silent — the model works, returns plausible output, and is simply miscalibrated for the population it serves\. The cost is not just performance but inclusion and representation: whether AI systems reflect and serve the societies that use them\[[3](https://arxiv.org/html/2608.02949#bib.bib21)\]\.

The two lanes are not independent\. The costs compound each other — and together they constitute aloss of sovereign capacityover a technology that is becoming critical infrastructure\[[17](https://arxiv.org/html/2608.02949#bib.bib14)\]\. A region that cannot build or evaluate AI for its own industries, languages, and institutions cannot capture the gains this technology offers, nor exercise meaningful agency over how it is used inside its borders\.

## 3State of Latin American datasets

The dataset layer faces two compounding problems that must be addressed together\. The first isdiscovery\[[14](https://arxiv.org/html/2608.02949#bib.bib33)\]; the second issupply\. They reinforce each other — which is why both must be in scope\.

Latin American AI datasets exist — the problem is that they are scattered across places like Hugging Face\[[20](https://arxiv.org/html/2608.02949#bib.bib28)\], the Mozilla Foundation, and the appendices of research papers, with no shared index\. A practitioner askingwhich datasets can I use to train a model for medical transcription in Rioplatense Spanish?can answer only by personal network or weeks of search\. This is not a failure of any platform; it is the absence of a regionally relevant index\.

Even with perfect indexing, the total volume of Latin American datasets would remain far below what frontier AI development requires — and far below what North America, Europe, and Asia have accumulated\. This is a structural gap, not a capability one\. The gap compounds annually — every year without an active creation effort is a year the distance grows\[[11](https://arxiv.org/html/2608.02949#bib.bib37)\]\.

These two problems reinforce each other in a loop that, left alone, would persist indefinitely\. There is no place to publish a Latin American AI dataset where the contributor sees a return on the act of contributing, so fewer datasets get published; because few are published, no critical mass forms to motivate the next contributor\[[21](https://arxiv.org/html/2608.02949#bib.bib38)\]\. The loop closes\. TheDataHubis designed to break it from both ends — indexing what already exists \(discovery\) and making contribution a visible, rewarded act \(supply\)\.

## 4A task\-first ontology for AI

Existing dataset catalogs almost universally organize entries by input modality or by target model — we believe both are wrong\. An AI model is fundamentally a program for a task, and task is therefore the first axis\. Tasks include transcription, classification, extraction, translation, forecasting, segmentation, and others\.

Existing dataset catalogs almost universally organize entries by input modality or by target model — we believe both are wrong for AI, because an AI model is fundamentally aprogramthat performs atask\[[1](https://arxiv.org/html/2608.02949#bib.bib39)\]\. This is also in line with how an AI practitioner works: he/she doesn’t start fromI have an imageorI want to use model X— they start fromI need a system that does Y\.\[[16](https://arxiv.org/html/2608.02949#bib.bib35)\]Thus the right ontology is one organized around what an AI systemdoes\(e\.g\. extraction, classification, reasoning, transcription, etc\.\)\. This DataHub adopts atask\-first ontologywith two nested levels:domain&language\. This is represented as/<task?\>/<domain?\>/<language?\>\.

Domain is the second axis because a task without domain context is almost always too broad to achieve the needed performance\. Within transcription, medical and legal are different datasets — vocabulary, annotation conventions, etc\. A model trained on medical transcription underperforms on legal\-domain audio files even within the same language\.

The third axis is language — and within language,regional variant\. Medical transcription in Rioplatense Spanish is not interchangeable with Colombian Spanish or Brazilian Portuguese\[[8](https://arxiv.org/html/2608.02949#bib.bib15)\]: clinical vocabulary, patient\-facing register, and acoustic patterns all differ\.

Each level compounds the value of the last: general transcription is useful, medical transcription is more useful, medical transcription in Rioplatense Spanish is more useful still\. The compounding pays off on both sides\. For industry, each layer of specificity buys a performance dividend — a Buenos Aires hospital deploying a medical\-transcription model tuned to Rioplatense Spanish gets a performance a general model cannot deliver\. For public institutions and the populations they serve, the same specificity buys representation and inclusivity — models that actually work for the people who use them\.

A practitioner walks this hierarchy from general to specific, and the ontology is designed so that walk is the natural one\. The same structure guides a contributor: declaring task, domain, and variant places a dataset in the exact slot where someone searching for it will look\. The ontology is the search interface and the publishing form in one\.

## 5Open by design, incentive\-driven by construction

On the geopolitical side, we prefer amultipolar AI world, and we believe Latin America should be one of those poles — with its own technology and its own agency\. This is a position against both single\-pole dominance — a world where AI is defined by the United States or China alone — and against a few\-pole world that excludes Latin America\. The more regions that develop their own AI systems, the better; this is the region’s stake in that outcome\. It is not an argument against collaboration with other regions; it is an argumentforLatin America contributing to that collaboration on its own terms\.

The question of how to build it — open or closed, shared or proprietary — is a strategic choice, not a neutral technical decision\. Different regions have taken different routes: some have built largely closed and proprietary AI ecosystems; others have built open and collaborative ones\[[5](https://arxiv.org/html/2608.02949#bib.bib40)\]\. Both can produce AI capability\. The choice shapes who can participate and how fast a non\-frontier region can close the gap\.

We argueopen by designis right for Latin America for a region\-specific reason: since collaboration fosters development, collaboration matters most for regions not at the frontier of a technology\[[9](https://arxiv.org/html/2608.02949#bib.bib24)\]\. For a non\-frontier region, the priority is not to protect what it has from competition but to contribute as much as possible, compounding collaborative efforts to close the gap\.

A commons is not self\-sustaining by virtue of being a commons\[[15](https://arxiv.org/html/2608.02949#bib.bib32)\]\. Without explicit incentives, there’s no intertia to continue publishing data into a hub\. Each new dataset compounds the value of every prior contribution — the index grows more complete, the ontology more populated\. Contributors gain something concrete: visibility within the regional AI community, attribution attached to their dataset, and recognition for their institution through brand\-awareness\.

The DataHub is not designed as a static catalog but as acontinuum: it has active contribution, curation, and growth\. The contribution path must be low\-friction — a researcher who has just produced a dataset should be able to publish it in minutes\[[4](https://arxiv.org/html/2608.02949#bib.bib41)\]\.

## 6Open problems

The DataHub does not resolve every question implied by the dataset layer it begins to build\.

Data protection, copyright, and public\-data access regimes differ country by country across Latin America\. The Hub cannot resolve this — these are matters of national law\. What it can and must do is surface the licensing and legal\-jurisdiction information of every indexed dataset\[[10](https://arxiv.org/html/2608.02949#bib.bib31)\], so practitioners decide with the relevant context visible\[[6](https://arxiv.org/html/2608.02949#bib.bib42)\]\.

There are no shared standards for dataset metadata, licensing, or quality in the region\[[7](https://arxiv.org/html/2608.02949#bib.bib30)\]\. Two datasets nominally aboutSpanish transcriptionmay use incompatible annotation schemas, incompatible licenses, and incompatible quality bars\[[13](https://arxiv.org/html/2608.02949#bib.bib43)\]Collaboration norms — attribution, versioning, and arbitration when a community contests how it is represented — also remain undefined\. When one party improves another’s dataset, how is credit allocated? These are questions for the regional community to decide together, not for one team to impose\.

Companies, universities, governments, and public bodies across Latin America hold significant volumes of data that would be valuable as published datasets — and most of it does not get published\. The right incentives for each of these to publish their data remain unknown\[[19](https://arxiv.org/html/2608.02949#bib.bib44)\]\. The DataHub’s visibility mechanisms address part of this — publishing now produces recognition\. But the deeper problem requires policy work outside the DataHub: grant requirements that mandate dataset publication, university promotion criteria that value dataset contributions, government open\-data policies that treat AI\-readable formats as a deliverable\.

## 7A kickstart and an invitation

This paper is akickstart— a working artifact and a position targeting the dataset layer specifically; the benchmark layer is the subject of LatamBoard, a companion effort\. The SURUS team built the first version of the Data Hub because somebody had to begin\. Together, the Data Hub and LatamBoard are an invitation — not a finished proposal — and not the work of a single team\.

To companies, researchers and institutions holding data: we invite you to publish it into the Hub\. Whether the dataset is a paper appendix from three years ago, an active working corpus, or an internal archive that has never been released — the DataHub is built to receive it\. Each contribution compounds the value of every prior one\.

To AI practitioners across the region: we invite you to use the DataHub and contribute to its improvement\. The most important signal that matters is practitioners choosing datasets from it, building models with it, and contributing datasets to it\. Usage and contribution is what converts a commons from a directory into infrastructure — and feedback from use is what improves it\.

To the broader community — universities, foundations, governments, companies: help shape what the foundational infrastructure for AI in Latin America should become, contribute, and then benefit from it\. The ontology will evolve, standards must be debated, and collaboration norms must be defined\. None of this is one team’s work — and none of the benefit accrues to a single organization\.

The dataset layer for Latin American AI will be built by the region, for the region, or it will not be built at all\. The DataHub is one step toward that future, and we invite you to join us\.

Availability\.The Data Hub is live at \[https://datahub\.lat\]\(https://datahub\.lat\)\.

## References

- \[1\]M\. Akhtaret al\.\(2024\)Croissant: a metadata format for ml\-ready datasets\.InProc\. Eighth Workshop on Data Management for End\-to\-End ML,External Links:[Document](https://dx.doi.org/10.1145/3650203.3663326)Cited by:[§4](https://arxiv.org/html/2608.02949#S4.p2.1)\.
- \[2\]S\. Balasubramaniamet al\.\(2021\)Machine learning based disease and pest detection in agricultural crops\.EAI Endorsed Transactions on Internet of Things\.External Links:[Document](https://dx.doi.org/10.4108/eetiot.5049)Cited by:[§2](https://arxiv.org/html/2608.02949#S2.p4.1)\.
- \[3\]E\. M\. Bender, T\. Gebru, A\. McMillan\-Major, and S\. Shmitchell\(2021\)On the dangers of stochastic parrots: can language models be too big?\.InProc\. 2021 ACM FAccT,External Links:[Document](https://dx.doi.org/10.1145/3442188.3445922)Cited by:[§2](https://arxiv.org/html/2608.02949#S2.p5.1)\.
- \[4\]C\. L\. Borgman\(2012\)The conundrum of sharing research data\.J\. of the American Society for Information Science and Technology\.External Links:[Document](https://dx.doi.org/10.1002/asi.22634)Cited by:[§5](https://arxiv.org/html/2608.02949#S5.p5.1)\.
- \[5\]D\. Corrales\-Garay, J\. L\. Rodríguez\-Sánchez, and A\. Montero\-Navarro\(2024\)Co\-creating value with ai: a bibliometric approach to the use of ai in open innovation ecosystems\.IEEE Access\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2024.3391054)Cited by:[§5](https://arxiv.org/html/2608.02949#S5.p2.1)\.
- \[6\]P\. N\. Edwards, M\. S\. Mayernik, A\. L\. Batcheller, G\. C\. Bowker, and C\. L\. Borgman\(2011\)Science friction: data, metadata, and collaboration\.Social Studies of Science\.External Links:[Document](https://dx.doi.org/10.1177/0306312711413314)Cited by:[§6](https://arxiv.org/html/2608.02949#S6.p2.1)\.
- \[7\]T\. Gebruet al\.\(2021\)Datasheets for datasets\.Communications of the ACM\.External Links:[Document](https://dx.doi.org/10.1145/3458723)Cited by:[§6](https://arxiv.org/html/2608.02949#S6.p3.1)\.
- \[8\]B\. Gonçalves and D\. Sánchez\(2015\)Learning about spanish dialects through twitter\.arXiv preprint arXiv:1511\.04970\.Cited by:[§4](https://arxiv.org/html/2608.02949#S4.p4.1)\.
- \[9\]J\. Lerner and J\. Tirole\(2002\)Some simple economics of open source\.The Journal of Industrial Economics50\(2\),pp\. 197–234\.External Links:[Document](https://dx.doi.org/10.1111/1467-6451.00174)Cited by:[§5](https://arxiv.org/html/2608.02949#S5.p3.1)\.
- \[10\]S\. Longpreet al\.\(2024\)A large\-scale audit of dataset licensing and attribution in ai\.Nature Machine Intelligence\.External Links:[Document](https://dx.doi.org/10.1038/s42256-024-00878-8)Cited by:[§6](https://arxiv.org/html/2608.02949#S6.p2.1)\.
- \[11\]A\. Mandal, S\. Leavy, and S\. Little\(2021\)Dataset diversity\.InProc\. 1st Intl Workshop on Trustworthy AI for Multimedia Computing,External Links:[Document](https://dx.doi.org/10.1145/3475731.3484956)Cited by:[§3](https://arxiv.org/html/2608.02949#S3.p3.1)\.
- \[12\]R\. Manvi, S\. Khanna, M\. Burke, D\. B\. Lobell, and S\. Ermon\(2024\)Large language models are geographically biased\.InICML 2024,Note:arXiv:2402\.02680Cited by:[§2](https://arxiv.org/html/2608.02949#S2.p2.1)\.
- \[13\]M\. A\. Musenet al\.\(2022\)Modeling community standards for metadata as templates makes data fair\.Scientific Data\.External Links:[Document](https://dx.doi.org/10.1038/s41597-022-01815-3)Cited by:[§6](https://arxiv.org/html/2608.02949#S6.p3.1)\.
- \[14\]F\. Nargesianet al\.\(2021\)Data lake management\.Proceedings of the VLDB Endowment\.External Links:[Document](https://dx.doi.org/10.14778/3352063.3352116)Cited by:[§3](https://arxiv.org/html/2608.02949#S3.p1.1)\.
- \[15\]E\. Ostrom\(2010\)Beyond markets and states: polycentric governance of complex economic systems\.American Economic Review100\(3\),pp\. 641–672\.External Links:[Document](https://dx.doi.org/10.1257/aer.100.3.641)Cited by:[§5](https://arxiv.org/html/2608.02949#S5.p4.1)\.
- \[16\]A\. Paullada, I\. D\. Raji, E\. M\. Bender, E\. Denton, and A\. Hanna\(2021\)Data and its \(dis\)contents: a survey of dataset development and use in machine learning research\.Patterns\.External Links:[Document](https://dx.doi.org/10.1016/j.patter.2021.100336)Cited by:[§4](https://arxiv.org/html/2608.02949#S4.p2.1)\.
- \[17\]H\. Roberts\(2024\)Digital sovereignty and artificial intelligence: a normative approach\.AI and Ethics\.External Links:[Document](https://dx.doi.org/10.1007/s10676-024-09810-5)Cited by:[§2](https://arxiv.org/html/2608.02949#S2.p6.1)\.
- \[18\]S\. Z\. Salas\-Pilco and Y\. Yang\(2022\)Artificial intelligence applications in latin american higher education: a systematic review\.Intl J\. of Educational Technology in Higher Education\.External Links:[Document](https://dx.doi.org/10.1186/s41239-022-00326-w)Cited by:[§1](https://arxiv.org/html/2608.02949#S1.p1.1)\.
- \[19\]C\. Tenopiret al\.\(2011\)Data sharing by scientists: practices and perceptions\.PLoS ONE\.External Links:[Document](https://dx.doi.org/10.1371/journal.pone.0021101)Cited by:[§6](https://arxiv.org/html/2608.02949#S6.p4.1)\.
- \[20\]T\. Wolfet al\.\(2019\)HuggingFace’s transformers: state\-of\-the\-art natural language processing\.arXiv preprint arXiv:1910\.03771\.Cited by:[§3](https://arxiv.org/html/2608.02949#S3.p2.1)\.
- \[21\]A\. Zuiderwijk, R\. Shinde, and W\. Jeng\(2020\)What drives and inhibits researchers to share and use open research data?\.PLOS ONE\.External Links:[Document](https://dx.doi.org/10.1371/journal.pone.0239283)Cited by:[§3](https://arxiv.org/html/2608.02949#S3.p4.1)\.

Similar Articles

On the missing benchmarks layer and a potential solution

arXiv cs.AI

This paper argues that Latin America lacks a benchmark layer for native AI development and proposes an open, task-first EvalsHub infrastructure, with LatamBoard as its first regional instance, to audit AI systems and direct optimization toward local needs.

The emergence of the web data infrastructure layer for AI

MIT Technology Review

This article discusses the growing need for a web data infrastructure layer to provide AI models with fresh, real-time, and trustworthy data, highlighting challenges like static training data and the importance of retrieval-augmented generation.

Open Source AI Gap Map (Website)

TLDR AI

The article introduces the Open Source AI Gap Map website, which evaluates over 24,626 projects to identify missing components in the open source AI stack, and invites collaborators to help close the gaps.

Everyone wants agents. Almost nobody has the data layer to run them.

Reddit r/AI_Agents

The article argues that many teams are eager to implement AI agents but overlook the foundational data layer, leading to fragmented integrations and maintenance debt. It emphasizes the importance of building a unified data retrieval system first to ensure scalability and efficiency in AI projects.

Voices in the Loop: Mapping Participatory AI

arXiv cs.AI

This paper presents a reproducible protocol for building an open repository and interactive atlas of participatory AI initiatives, analyzing 131 records to reveal geographic and lifecycle patterns, and proposes a framework for participatory-by-default AI infrastructures.