迈向「模型即库」:面向低资源非洲语言的离线、社区众包式 AI

arXiv cs.AI 论文

摘要

这篇立场论文提出了「模型即库」(Model as a Library,简称 MaaL)这一软件架构,它将小型、由社区招募数据训练而成的语音模型打包为带版本管理的设备端依赖项,用于低资源非洲语言场景下离线且无幻觉的数据采集。论文提出利用关键词识别,将封闭词表的数字表单转变为面向识字率较低人群的语音化、设备端数据采集流程,并将现有的表单工具转译为 MaaL 模式。

arXiv:2609.38574v1 Announce Type: new Abstract: Large language models are frequently proposed as a route to AI-powered services for African communities, but they are least reliable exactly where the need is greatest: all African languages remain low-resource by any standard measure, and models trained on scraped, standardised text systematically misrepresent the dialectal and regional variation of how people actually speak. We introduce \textbf{Model as a Library (MaaL)}, a software architecture that packages small, community-enrolled speech models as versioned on-device dependencies, enabling offline structured data collection that cannot generatively hallucinate, for populations that current language models serve worst. Rather than relying on web-scraped corpora, MaaL's vocabulary is enrolled directly from a small number of example recordings by the speakers themselves, at the point of deployment. We describe the architecture and its central mechanism - keyword spotting that turns a closed-vocabulary text form into a voice form, filled and submitted entirely on-device - and propose transpiling the closed-vocabulary elements already present in widely-deployed digital form tools into MaaL schemas, a low-friction path to voice-first, offline data collection for the low-literacy populations these tools already reach. This is a position and system-design paper: we describe the concept, the mechanism, and an analytical feasibility case, and identify what a working implementation still requires.
查看原文
查看缓存全文

缓存时间: 2026/10/01 09:41

# Towards Model as a Library: Offline, Community-Sourced AI for Low-Resource African Languages
Source: [https://arxiv.org/html/2609.38574](https://arxiv.org/html/2609.38574)
\\workshoptitle

GlobalSouthAI

Fendji K\. E\. Jean LouisAffiliation:Centre of Research, Experimentation and ProductionAffiliation:SCEMI, University of Ngaoundere, Ngaoundere, CameroonAffiliation:Stellenbosch Institute for Advanced StudyAffiliation:Wallenberg Research Centre at Stellenbosch University, Stellenbosch, South AfricaEmail:[jl\.fendji@egcim\-univ\-ndere\.cm](mailto:)

###### Abstract

Large language models are frequently proposed as a route to AI\-powered services for African communities, but they are least reliable exactly where the need is greatest: all African languages remain low\-resource by any standard measure, and models trained on scraped, standardised text systematically misrepresent the dialectal and regional variation of how people actually speak\. We introduceModel as a Library \(MaaL\), a software architecture that packages small, community\-enrolled speech models as versioned on\-device dependencies, enabling offline structured data collection that cannot generatively hallucinate, for populations that current language models serve worst\. Rather than relying on web\-scraped corpora, MaaL’s vocabulary is enrolled directly from a small number of example recordings by the speakers themselves, at the point of deployment\. We describe the architecture and its central mechanism – keyword spotting that turns a closed\-vocabulary text form into a voice form, filled and submitted entirely on\-device – and propose transpiling the closed\-vocabulary elements already present in widely\-deployed digital form tools into MaaL schemas, a low\-friction path to voice\-first, offline data collection for the low\-literacy populations these tools already reach\. This is a position and system\-design paper: we describe the concept, the mechanism, and an analytical feasibility case, and identify what a working implementation still requires\.

Keywords:offline AI, keyword spotting, low\-resource languages, on\-device inference, few\-shot learning, human\-computer interaction

## 1Introduction

Large language models are increasingly proposed as infrastructure for education, health, and agricultural services across the Global South\. Yet the evidence on their reliability for African languages is not encouraging\. African languages remain low\-resource by any standard measure, and lower resource levels directly translate into lower\-quality, less reliable models – an inversion in which the populations with the most to gain are served worst\. The failure mode is not abstract: on the WARRI benchmark for West African Pidgin, models adapted to the standardised media register \(BBC variety\) score 76\.3–83\.4 ChrF\+\+, but the same models drop to approximately 54 on everyday community\-level Naija – a 24\.3\-point gap that causes model reasoning to drift from accuracy toward linguistic guesswork whenever input departs from the training distribution\[[1](https://arxiv.org/html/2609.38574#bib.bib9),[5](https://arxiv.org/html/2609.38574#bib.bib1)\]\. This "Standard English bias" is driven by training data that overrepresents one register and erases the rest\. Machine\-translation pipelines, often proposed as a workaround, introduce their own compounding and poorly understood errors\. The underlying cause is not a prompting problem; it is data scarcity, and no amount of scale fixes a corpus that was never representative to begin with\.

This is a structural, not incidental, problem for deployment\. Large parts of the population most in need of these tools also face the least reliable connectivity and linguistic coverage simultaneously – so even where a capable multilingual model exists, connectivity and language together block access, widening the very data gap that caused the problem\[[4](https://arxiv.org/html/2609.38574#bib.bib13)\]\. A tool requiring a live connection to a distant server is, for this population, not a language\-access solution but a second barrier stacked on the first\.

We argue that for a large and important class of applications – structured data collection, not open\-ended dialogue – neither of these failure modes is necessary\. We introduceModel as a Library \(MaaL\), a system architecture in which trained models are packaged as versioned, on\-device software dependencies, analogous to code libraries, rather than accessed as a remote service\. A MaaL*domain model*is a small keyword\-spotting classifier whose vocabulary is enrolled directly from a small number of example recordings – typically 5–10 in published few\-shot keyword\-spotting work\[[8](https://arxiv.org/html/2609.38574#bib.bib10),[9](https://arxiv.org/html/2609.38574#bib.bib11)\]– by the speakers themselves, collected on\-device, with no cloud round\-trip and no retraining\. Because inference is closed\-set matching rather than open\-ended generation, MaaL cannot hallucinate a response in the way a generative model can: it returns either a term the community itself provided, or an explicit rejection\.

We make three contributions\. First, we describe the MaaL architecture and its core primitive, the*model\-typed input*, and argue for its specific relevance to the African\-language data problem\. Second, we work through a representative structured data\-collection scenario – a closed\-vocabulary agricultural survey – to make the mechanism concrete, and report analytical feasibility estimates for the resulting system\. Third, we propose mechanically transpiling the closed\-vocabulary elements already present in existing digital form tools \(ODK, KoboToolbox\) into MaaL schemas, giving low\-literacy respondents a path to complete forms unassisted, in their own language, without requiring organisations to build new data\-collection infrastructure from scratch\.

## 2Related Work

Model as a Service \(MaaS\), the dominant paradigm for accessing large models via cloud APIs, is structurally unsuited to intermittent\-connectivity deployment\[[12](https://arxiv.org/html/2609.38574#bib.bib2)\]\.Community\-trained small language modelsare the closest precedent: InkubaLM, a 400M\-parameter model trained from curated data across five African languages, matches or exceeds much larger models on in\-language tasks\[[11](https://arxiv.org/html/2609.38574#bib.bib3)\], validating that deliberately sourced small models beat scraped large ones for underrepresented languages\.Edge AI platformssuch as Edge Impulse package multiple models into a deployable firmware image\[[7](https://arxiv.org/html/2609.38574#bib.bib4)\], but require cloud retraining and full reflash for every vocabulary change; MaaL’s on\-device enrolment removes this dependency\.Mobile data collection platforms\(ODK, CommCare\) solve offline form synchronisation but leave AI inference cloud\-side\[[6](https://arxiv.org/html/2609.38574#bib.bib5)\]; MaaL is the missing offline inference layer for exactly these platforms\.Few\-shot and query\-by\-example keyword spottingprovides the closed\-set matching mechanism MaaL depends on: prototypical\-network approaches enrol new keywords from a handful of examples without retraining\[[8](https://arxiv.org/html/2609.38574#bib.bib10)\], open\-set variants explicitly separate enrolled terms from out\-of\-vocabulary input\[[9](https://arxiv.org/html/2609.38574#bib.bib11)\], and on\-device domain adaptation has been demonstrated under sub\-10KB RAM budgets\[[3](https://arxiv.org/html/2609.38574#bib.bib12)\]– MaaL packages this class of method as a versioned software dependency rather than proposing a new one\.

## 3The MaaL Architecture

A MaaL domain model is a tupleM=\(B,P,L,τ,m\)M=\(B,P,L,\\tau,m\), whereBBis a frozen embedding backbone mapping input of modalitymmto a fixed\-dimensional vector space;PPis a prototype store of enrolled keyword embeddings;LLis a label map from keywords to normalised output values; andτ\\tauis a confidence threshold below which the model returns an explicit rejection rather than a guess\. Inference is nearest\-prototype matching: embed the input, compare by cosine similarity against every enrolled prototype, accept if the best match clearsτ\\tau\. This is a closed\-set operation by construction – the model can only return something a community member actually said, or nothing at all\. This closed\-set property eliminates generative fabrication – MaaL cannot invent a term no one said – but it does not eliminate misrecognition: a wrong enrolled term can clearτ\\tau\(false accept\), and a correctly\-spoken term can fall short of it \(false reject\)\. These are classification errors, not hallucinations, but in health and agricultural data collection they carry real harm and are not yet analysed here\.

The system has three layers \(Table[1](https://arxiv.org/html/2609.38574#S3.T1)\)\. A*form schema*declares, for each field, which domain model applies – the only artefact a developer writes\. A*nomadic runtime*plays the prompt, records the response, routes it to the declared model, and applies the threshold, entirely on\-device\. The*model library*holds one shared embedding backbone \(trained once, centrally, on general multilingual speech data\) plus many small, independently versioned domain models, each contributing only a prototype store of roughly 512 bytes per enrolled term\.

Table 1:The three\-layer MaaL architecture\.Enrolling a new term requires a small number of example recordings passed once through the frozen backbone; the new prototype is their mean embedding\. No gradient computation, GPU, or connectivity is required, so a field worker can add a locally\-specific term – a crop variety, a market name – in minutes, directly from the language as spoken by the people who will use the system\.

## 4Illustration: From a Text Form to a Voice Form

To make the mechanism concrete, consider a representative closed\-vocabulary structured survey of the kind widely used in agricultural extension and monitoring programmes, with questions such as: which crop was planted, how many units of seed were used, in which season planting occurred, and whether a given input was applied\. Each question maps onto one of a small number of reusable domain\-model types:agri\_term\(a bounded set of crop names, mapped to normalised identifiers\),numeric\(spoken quantities\),time\_period\(seasons or months\), andyes\_no\(affirmatives and negatives\)\.numericis architecturally distinct from the other three types: spoken quantities are compositional rather than a small closed set, and nearest\-prototype matching does not by itself explain how arbitrary multi\-digit numbers are handled\. A practical approach – bounding the field to a discretised range of enrolled values, or composing digit\-level matches – is left unspecified here and is a concrete design question, not yet a solved one\. These four types cover a large share of the fields in a typical structured survey and, once built, are reusable across many forms and deployment languages\.

For any given deployment language, the vocabulary for each model – the specific crop names, season terms, and affirmative/negative pairs actually spoken – is compiled from community\-verifiable sources and confirmed with speakers directly, rather than scraped from web text: the same design choice that makes MaaL resistant to the dialect\-flattening failure mode described in Section 1\. A term that cannot be confirmed against a reliable source should be left marked for verification rather than guessed, a discipline a generative model has no equivalent mechanism to enforce\. At runtime, a spoken response to each question is embedded and matched against the enrolled vocabulary for that field’s domain model; an accepted match writes a normalised value into the corresponding form field, and the completed record is assembled and submitted exactly as a text\-form submission would be, without any field ever having been read or typed\.

### 4\.1From Text Forms to Audio\-First Forms

A practical obstacle to adopting any new data\-collection paradigm is that organisations across the Global South already have large investments in existing digital form tools – ODK, KoboToolbox, and similar XForm\-based platforms\[[6](https://arxiv.org/html/2609.38574#bib.bib5)\]are in active use across health, agriculture, and humanitarian programmes, but their interfaces assume the respondent, or an intermediary, can read\.

Closed\-vocabulary form elements – an XForm<select1\>, a dropdown, an HTML<select\>– already enumerate exactly the bounded vocabulary a MaaL domain model needs: the option list*is*a prototype\-store vocabulary, and the declared answer values*are*a label mapLL\. A numeric field binds directly to anumeric\-type domain model; a yes/no radio group binds toyes\_no\. A large class of existing text forms can therefore be*mechanically transpiled*into a MaaL form schema \(Table[1](https://arxiv.org/html/2609.38574#S3.T1), Layer 1\): parse the existing form definition, map each closed\-vocabulary field to a domain model seeded with its option list, and enrol prototypes for that vocabulary from community recordings, as illustrated above\. The organisation’s existing form logic and question ordering require no change\.

This reframes MaaL as an accessibility layer underneath data\-collection infrastructure already deployed across the region, rather than a replacement for it – turning a form previously completable only with a literate intermediary reading questions aloud into one a respondent can complete unassisted, in their own language, offline\. Building and evaluating such a transpiler is a concrete near\-term extension of this work\.

## 5Feasibility

We report analytical estimates grounded in published DS\-CNN benchmarks, not measurements from a deployed system \(Table[2](https://arxiv.org/html/2609.38574#S5.T2)\)\. These estimates cover on\-device storage and latency only; they say nothing about recognition accuracy, which depends on few\-shot enrolment quality under real acoustic conditions and is not addressed by parameter counts or benchmark footprints\. Implementing and evaluating the real backbone on field\-collected audio in a target deployment language is the paper’s central open question\. The DS\-CNN family\[[13](https://arxiv.org/html/2609.38574#bib.bib6)\]spans roughly 39K \(DS\-CNN\-S\) to 189K \(DS\-CNN\-M\) parameters; after INT8 quantisation, DS\-CNN\-S measures 52\.5 KB in the MLPerf Tiny reference benchmark\[[2](https://arxiv.org/html/2609.38574#bib.bib7)\]and 46\.15 KB in an independent hardware measurement on a Raspberry Pi Pico 2\[[10](https://arxiv.org/html/2609.38574#bib.bib8)\]– two sources converging on∼\\sim46–52 KB for the smaller variant\. Scaling this ratio to DS\-CNN\-M gives an approximate upper bound of∼\\sim190 KB, an extrapolation rather than a third direct measurement\. The resulting footprint and latency figures \(Table[2](https://arxiv.org/html/2609.38574#S5.T2)\) are comfortably within the storage and compute budget of commodity Android devices and microcontroller\-class hardware, and compatible with the natural pacing of a spoken interaction\.

Table 2:Analytical feasibility estimates \(literature\-grounded, not measured\)\.
## 6Discussion and Open Questions

MaaL’s narrowness is deliberate: it cannot answer an open question or handle a request outside its enrolled vocabulary, and by design it should not attempt to\. Where richer capability is genuinely needed, we favour an explicit, visually distinct escalation to a community\-trained generative model such as InkubaLM\[[11](https://arxiv.org/html/2609.38574#bib.bib3)\], rather than blurring the boundary that gives MaaL its zero\-hallucination property\. Open questions include: accuracy of enrolled models under real field acoustic conditions rather than synthetic audio, including the false\-accept/false\-reject tradeoff as a function ofτ\\tauand its associated misrecognition harms; the minimum viable shared backbone across a cluster of related regional languages sharing a deployment area; robustness of a text\-to\-audio\-form transpiler \(Section 4\.1\) across heterogeneous existing schemas; and, most centrally for this venue, what governance structure should oversee community vocabulary curation itself – who decides what a term means, who may extend it, and how disagreement about a translation or dialect boundary is resolved, rather than settled implicitly by whoever built the tool\.

## 7Conclusion

Large language models are least reliable exactly where linguistic need is greatest, and the cause is a data gap that scale alone cannot fix\. MaaL proposes a narrower but more honest alternative: small, community\-enrolled, closed\-set models that cannot misrepresent what was not said, work fully offline, and grow their vocabulary from the language as it is actually spoken rather than from a scraped corpus\. Treating the closed\-vocabulary structure already present in widely\-deployed digital form tools as a ready\-made source of that vocabulary also offers a route to adoption that does not require organisations to discard existing infrastructure\.

## Acknowledgments and Disclosure of Funding

This work received no dedicated research grant funding from any agency in the public, commercial, or not\-for\-profit sectors\. The author gratefully acknowledges the Stellenbosch Institute for Advanced Study \(STIAS\), Stellenbosch, South Africa, for the fellowship and residency during which this work was developed\.

We used Claude \(Anthropic\) to assist with drafting and editing this manuscript, deriving the analytical estimates in Table[2](https://arxiv.org/html/2609.38574#S5.T2)from cited literature values, and researching and verifying citations\. All citations, technical claims, and derivations were independently checked against primary sources before submission\. The MaaL architecture, the model\-typed input primitive, and the audio\-first\-forms proposal originate from the author; no LLM is a component of the proposed system\.

## References

- \[1\]\(2025\)Does generative AI speak Nigerian\-Pidgin?: issues about representativeness and bias for multilingualism in LLMs\.InFindings of the Association for Computational Linguistics: NAACL 2025,pp\. 1571–1583\.Cited by:[§1](https://arxiv.org/html/2609.38574#S1.p1.1)\.
- \[2\]C\. Banbury, V\. J\. Reddi, P\. Torelli, J\. Holleman, N\. Jeffries, C\. Király, P\. Montino, D\. Kanter, S\. Ahmed, D\. Pau, U\. Thakker, A\. Torrini, P\. Warden, J\. Cordaro, G\. Di Guglielmo, J\. Duarte, S\. Gibellini, V\. Parekh, H\. Tran, N\. Tran, N\. Wenxu, and X\. Xuesong\(2021\)MLPerf tiny benchmark\.Proceedings of Machine Learning and Systems \(MLSys\) / NeurIPS Datasets and Benchmarks Track, arXiv:2106\.07597\.Cited by:[Table 2](https://arxiv.org/html/2609.38574#S5.T2.4.2.3.1.1),[§5](https://arxiv.org/html/2609.38574#S5.p1.1)\.
- \[3\]C\. Cioflan, L\. Cavigelli, M\. Rusci, M\. de Prado, and L\. Benini\(2024\)On\-device domain learning for keyword spotting on low\-power extreme edge systems\.arXiv:2403\.10549\.Cited by:[§2](https://arxiv.org/html/2609.38574#S2.p1.1)\.
- \[4\]J\. L\. K\. E\. Fendji\(2024\)From left behind to left out: generative AI or the next pain of the unconnected\.Harvard Data Science Review\.Note:Special Issue 5External Links:[Document](https://dx.doi.org/10.1162/99608f92.427c5ab9)Cited by:[§1](https://arxiv.org/html/2609.38574#S1.p2.1)\.
- \[5\]A\. O\. Gabriel and A\. R\. Yusuf\(2026\)Adversarial fragility and language vulnerability in clinical ai: a systematic audit of diagnostic collapse under imperceptible perturbations and cross\-lingual drift in low\-resource healthcare settings\.arXiv preprint arXiv:2605\.16993\.Cited by:[§1](https://arxiv.org/html/2609.38574#S1.p1.1)\.
- \[6\]C\. Hartunget al\.\(2010\)Open data kit: tools to build information services for developing regions\.InProceedings of ACM ICTD,Cited by:[§2](https://arxiv.org/html/2609.38574#S2.p1.1),[§4\.1](https://arxiv.org/html/2609.38574#S4.SS1.p1.1)\.
- \[7\]S\. Hymelet al\.\(2022\)Edge impulse: an mlops platform for tiny machine learning\.arXiv preprint arXiv:2212\.03332\.Cited by:[§2](https://arxiv.org/html/2609.38574#S2.p1.1),[Table 2](https://arxiv.org/html/2609.38574#S5.T2.4.4.3.1.1)\.
- \[8\]A\. Parnami and M\. Lee\(2022\)Few\-shot keyword spotting with prototypical networks\.InProc\. ICMLT,Cited by:[§1](https://arxiv.org/html/2609.38574#S1.p3.1),[§2](https://arxiv.org/html/2609.38574#S2.p1.1)\.
- \[9\]M\. Rusci and T\. Tuytelaars\(2023\)Few\-shot open\-set learning for on\-device keyword spotting\.arXiv:2306\.02161\.Cited by:[§1](https://arxiv.org/html/2609.38574#S1.p3.1),[§2](https://arxiv.org/html/2609.38574#S2.p1.1)\.
- \[10\]Ł\. Sobczak and K\. Fonał\(2026\)Efficient Polish\-language keyword spotting on microcontrollers: compact neural architectures, quantization, and on\-device validation on the Raspberry Pi Pico 2\.Applied Sciences16\(15\),pp\. 7844\.External Links:[Document](https://dx.doi.org/10.3390/app16157844)Cited by:[Table 2](https://arxiv.org/html/2609.38574#S5.T2.4.2.3.1.1),[§5](https://arxiv.org/html/2609.38574#S5.p1.1)\.
- \[11\]A\. L\. Tonjaet al\.\(2024\)InkubaLM: a small language model for low\-resource african languages\.arXiv preprint\.Cited by:[§2](https://arxiv.org/html/2609.38574#S2.p1.1),[§6](https://arxiv.org/html/2609.38574#S6.p1.1)\.
- \[12\]Z\. Wenet al\.\(2022\)Model\-as\-a\-service: a survey\.arXiv preprint arXiv:2211\.01078\.Cited by:[§2](https://arxiv.org/html/2609.38574#S2.p1.1)\.
- \[13\]Y\. Zhang, N\. Suda, L\. Lai, and V\. Chandra\(2017\)Hello edge: keyword spotting on microcontrollers\.arXiv preprint arXiv:1711\.07128\.Cited by:[Table 2](https://arxiv.org/html/2609.38574#S5.T2.4.2.3.1.1),[Table 2](https://arxiv.org/html/2609.38574#S5.T2.4.3.3.1.1),[§5](https://arxiv.org/html/2609.38574#S5.p1.1)\.

相似文章

大语言模型在低资源语言人文学科研究中的机遇与挑战

arXiv cs.CL

本文系统评估了大语言模型在低资源语言研究中的应用,分析了在语言变异、历史文献、文化表达和文学分析等方面的机遇与挑战。研究强调了跨学科合作和定制化模型开发,以保护语言和文化遗产,同时解决数据可获取性、模型适应性和文化敏感性问题。

用于空中作战决策支持的离线多模态大型语言模型

arXiv cs.AI

本文研究离线多模态大型语言模型作为空中作战的决策支持工具,详细描述了一种模块化检索增强架构,并介绍了与Brazilian Air Force合作的试点研究,该研究显示在条令评估任务中认知负荷降低且效率提高。

微软测试新的 MAI Realtime 语音模型(2分钟阅读)

TLDR AI

微软正在其 MAI Playground 上以早期访问形式测试一款新的原生实时语音模型 MAI Realtime。该全双工系统支持多种语言、低延迟和可配置的轮流对话,定位为 OpenAI GPT Live 和 Sesame 的竞争对手。