Mwando: Leveraging AI to Preserve and Teach shiKomori

arXiv cs.CL Papers

Summary

Mwando is a virtual educational assistant leveraging AI to preserve and teach the Comorian language shiKomori, utilizing a multi-agent architecture with vector search, knowledge graph, and web fallback.

arXiv:2607.23481v1 Announce Type: new Abstract: This paper presents Mwando, a virtual educational assistant designed to support the teaching and preservation of shiKomori, the language of the Comoros Islands. The system covers the four main dialectal variants (shiNgazidja, shiMwali, shiNdzuani and shiMaore) through a knowledge base constructed from phrases, proverbs, dictionaries and grammar lessons. A multi-agent architecture combining vector search, a knowledge graph and web search fallback enables accurate and context-aware responses. Evaluation on 500 queries demonstrates strong performance on vocabulary lookup and grammar explanations, while qualitative case studies illustrate both capabilities and current limitations. This work represents an initial step toward computational support for shiKomori and provides a blueprint for developing AI-powered educational tools for other low-resource languages.
Original Article
View Cached Full Text

Cached at: 07/28/26, 06:29 AM

# Mwando: Leveraging AI to Preserve and Teach shiKomori
Source: [https://arxiv.org/html/2607.23481](https://arxiv.org/html/2607.23481)
Naira Abdou Mohamed1, Haidar Nassur Said Ali2, Mohamed Hazra3, Naoufal Mohamed Soibira1,4,Roushnaty Ali Yamani3 1Rifai, Moroni, Comoros 2CP2BM Lab \- Hassan II University, Casablanca, Morocco 3Ibn Tofail University, Kenitra, Morocco 4Sciences Po Grenoble, Grenoble, France

###### Abstract

This paper presentsMwando, a virtual educational assistant designed to support the teaching and preservation of shiKomori, the language of the Comoros Islands\. The system covers the four main dialectal variants \(shiNgazidja, shiMwali, shiNdzuani and shiMaore\) through a knowledge base constructed from phrases, proverbs, dictionaries and grammar lessons\. A multi\-agent architecture combining vector search, a knowledge graph and web search fallback enables accurate and context\-aware responses\. Evaluation on 500 queries demonstrates strong performance on vocabulary lookup and grammar explanations, while qualitative case studies illustrate both capabilities and current limitations\. This work represents an initial step toward computational support for shiKomori and provides a blueprint for developing AI\-powered educational tools for other low\-resource languages\.

Mwando: Leveraging AI to Preserve and Teach shiKomori

## 1Introduction

Language preservation is a central concern for international organizations such as the United Nations Development Programme \(UNDP\), which emphasizes African languages as key drivers of the continent developmentUnited Nations Development Programme and Ministry of Enterprises and Made in Italy \([2024](https://arxiv.org/html/2607.23481#bib.bib28)\)\. However, this potential remains largely unrealized, as very few technological solutions currently support these languagesWild \([2025](https://arxiv.org/html/2607.23481#bib.bib29)\)\. This situation is problematic, as language and culture are major drivers of developmentRotondo \([2016](https://arxiv.org/html/2607.23481#bib.bib32)\)and play a central role in everyday life across African societies\. However, advances in language processing technologies, particularly generative AI systems, offer a promising opportunity: when appropriately adapted to cultural and linguistic specificities, these models can serve as effective conduits for the preservation of cultural heritageColaceet al\.\([2025](https://arxiv.org/html/2607.23481#bib.bib30)\); Koc \([2025](https://arxiv.org/html/2607.23481#bib.bib31)\)\.

In this work, we aim to contribute to the initial development of Natural Language Processing \(NLP\) technologies for shiKomori, the language spoken in the Comoros Islands\. Building on our previous work on foundational NLP resources for Comorian and its dialectal variationsNairaet al\.\([2024](https://arxiv.org/html/2607.23481#bib.bib27),[2025](https://arxiv.org/html/2607.23481#bib.bib22)\), we shift the focus to an applied educational setting\.

Specifically, we introduce Mwando, a virtual assistant designed for teaching shiKomori, with the goal of supporting the transmission and preservation of Comorian culture\. Mwando \("beginning" in shiKomori\) honors Sheikh Ahmed Kamar\-Eddine’s pioneering work on Comorian language documentation and writing standardizationLafon \([2007](https://arxiv.org/html/2607.23481#bib.bib21)\); Nairaet al\.\([2025](https://arxiv.org/html/2607.23481#bib.bib22)\)\. Our contributions are threefold:

- •We compile and organize a multi\-source corpus for shiKomori, covering phrases, proverbs, dictionaries and grammar lessons and make these data publicly available to support future research\.
- •We develop a virtual assistant for educational purposes, capable of interacting with learners in shiKomori\.
- •We evaluate the assistant’s potential for supporting language learning and cultural preservation, providing insights for further development of NLP resources for extremely low\-resource languages\.

## 2Motivations

The arrival of generative AI has been a serious game changer in educationTeamet al\.\([2025](https://arxiv.org/html/2607.23481#bib.bib6)\)and it was immediately adopted by students and language teachersZaimet al\.\([2025](https://arxiv.org/html/2607.23481#bib.bib13)\)\. Among the reasons justifying this rapid adoption is the possibility for students to obtain quick feedback through a virtual teacher and for teachers, the ability to quickly and effectively customize lessons according to different casesGalaczi and Luckin \([2024](https://arxiv.org/html/2607.23481#bib.bib14)\)\. It is precisely along these lines that a project was adopted in Senegal in elementary education to support French teachers in better teaching in classrooms composed of students with significant cultural and linguistic diversityAgence Française de Développement \([2025](https://arxiv.org/html/2607.23481#bib.bib12)\)\.

In the Comoros, although the use of AI and information technology in education is very poorly documentedRoukiyat \([2026](https://arxiv.org/html/2607.23481#bib.bib9)\), there are previous studies that have emphasized the importance of taking shiKomori into account in teaching for greater effectivenessDANIEL \([2024](https://arxiv.org/html/2607.23481#bib.bib11)\)\. And local initiatives such as the Maecha association have set themselves the goal of literacy in Comorian for the adult population, which is composed of approximately 49\.7% illiterate individualsChauvet \([2015](https://arxiv.org/html/2607.23481#bib.bib10)\)\.

Consequently, it is crucial to consider modern solutions for learning shiKomori\. Such initiatives could not only benefit other categories of people and sectors, such as improving the experience of tourists in the archipelago, but they would also play an important role in the country’s sustainable developmentRoukiyat \([2026](https://arxiv.org/html/2607.23481#bib.bib9)\)\. Furthermore, they would hold strong potential for the inclusion of descendants of the Comorian diaspora, particularly those in France, whose population is estimated at more than 300,000 peopleLe Monde \([2024](https://arxiv.org/html/2607.23481#bib.bib7)\)and who contribute significantly to the development of the archipelagoAbdillahi \([2012](https://arxiv.org/html/2607.23481#bib.bib8)\)\.

## 3Linguistic Resources

One of the biggest challenges for this research was obtaining the data\. Our approach was to consult every possible resource available in order to build our knowledge base\. However, to ensure the reliability of the dataset, manual checks were necessary at each stage of processing\. In this section, we therefore describe all the corpora used, from the raw data to the transformed data for this work\.

### 3\.1Language Overview

Spoken exclusively in the Comoros archipelago, shiKomori is a Bantu language, closely related to SwahiliAhmed Chamanga \([2022](https://arxiv.org/html/2607.23481#bib.bib25)\)and to the Sabaki language groupServa and Pasquini \([2021](https://arxiv.org/html/2607.23481#bib.bib24)\)\. It consists of four dialectal varieties, each primarily associated with one island, yet exhibiting a high degree of mutual intelligibility among speakersAhmed Chamanga \([2022](https://arxiv.org/html/2607.23481#bib.bib25)\); Nairaet al\.\([2024](https://arxiv.org/html/2607.23481#bib.bib27)\)\. ShiKomori holds the status of a national language in the archipelago, while French and Arabic are the official languages\. In practice, shiKomori is mainly used orally in everyday communication and informal media, French dominates administrative and formal educational domains and Arabic is predominantly used in religious contextsChauvet \([2015](https://arxiv.org/html/2607.23481#bib.bib10)\)\.

In written form, shiKomori can be transcribed using either the Latin alphabet or the Arabic scriptLafon \([2007](https://arxiv.org/html/2607.23481#bib.bib21)\); Nairaet al\.\([2025](https://arxiv.org/html/2607.23481#bib.bib22)\)\. However, its written use remains limited and is characterized by a lack of orthographic standardization\. Despite recent initiatives aimed at improving the representation of African languages in NLPAdebaraet al\.\([2025](https://arxiv.org/html/2607.23481#bib.bib23)\), shiKomori remains largely underrepresented in the field, except for a few prior worksNairaet al\.\([2024](https://arxiv.org/html/2607.23481#bib.bib27),[2025](https://arxiv.org/html/2607.23481#bib.bib22)\); Abdourahamaneet al\.\([2016](https://arxiv.org/html/2607.23481#bib.bib26)\)\.

### 3\.2Data Sources

This work relies on six primary data sources, illustrated in Figure[1](https://arxiv.org/html/2607.23481#S3.F1)\. These sources are grouped into four main categories of data:

- •Useful phrases: These consist of commonly used phrases in shiKomori, without distinction between dialectal varieties\. The phrases are translated into English and cover a wide range of topics, from simple greetings to expressions commonly used in contexts such as markets, travel and everyday interactions\. The data were obtained from a Google Drive repository gathering various resources related to the Comoros Islands111[https://drive\.google\.com/drive/folders/17C\_03qCMm2rGDhgGqi\_8PKJfR30V\_S96](https://drive.google.com/drive/folders/17C_03qCMm2rGDhgGqi_8PKJfR30V_S96)\.
- •Proverbs: The proverbs were collected from the Instagram page*ProverbesComoriens*222[https://www\.instagram\.com/proverbescomoriens/](https://www.instagram.com/proverbescomoriens/)using the Apify scraping tool\. Each post contains proverbs in shiKomori along with their French and English translations\. However, the posts are published as images\. To extract the textual content, we used DeepSeek\-OCRWeiet al\.\([2025](https://arxiv.org/html/2607.23481#bib.bib20)\)\. This OCR\-based extraction occasionally introduced minor errors, particularly with diacritics and special characters, which were manually corrected during the data cleaning phase\.
- •Dictionaries: Two dictionaries were used\. The first, focusing on shiNgazidja, is a notable work produced by the Bahari Foundation333[https://fr\.scribd\.com/document/619345660/ShiNgazidja\-English\-Dictionary](https://fr.scribd.com/document/619345660/ShiNgazidja-English-Dictionary), which provides shiNgazidja words and expressions translated into English, with explicit annotation of nominal and grammatical classes\. The second dictionary was obtained from the same Google Drive repository mentioned above and contains shiMwali words translated into English, along with grammatical class annotations\.
- •

![Refer to caption](https://arxiv.org/html/2607.23481v1/x1.png)Figure 1:Data Preparation Pipeline\.Finally, all data sources used in this work are publicly available online\. For the Instagram data, we only collected publicly accessible posts and complied with the platform’s terms of service\. The dictionaries and grammar manuals are freely distributed by their respective publishers or hosting platforms\. We release our processed dataset under an open license to facilitate future research on Comorian language processing \(see Table[1](https://arxiv.org/html/2607.23481#S3.T1)\)\.

Table 1:Overview of the datasets used in this work\.
### 3\.3Knowledge Base Construction

A processing step was necessary to better structure the data\. Indeed, among the collected data, for example, in the case of grammar lessons, some data were in the form of tables, as illustrated in Figure[1](https://arxiv.org/html/2607.23481#S3.F1)\. Although Large Language Models \(LLM\) are effective at understanding textual data, they exhibit limitations when processing text mixed with tabular contentLiuet al\.\([2024](https://arxiv.org/html/2607.23481#bib.bib19)\)\. This issue primarily arises during the data tokenization phase\. To mitigate this, we first applied a document structuring step by converting the documents into Markdown format using the Word2md platform666[https://word2md\.com/](https://word2md.com/)\.

Subsequently, the Markdown content was fed into the ChatGPT API to restructure the data into a more easily exploitable knowledge base\. This process involved extracting elements such as chapter sections, covered concepts, tables, definitions and related content and organizing them into JSON files\. The specific structuring strategy varied depending on the type of data\. In all cases, a manual validation step was conducted to review and correct the outputs generated by ChatGPT\.

## 4Chatbot Workflow

In recent years, the use of multi\-agent approaches has significantly improved the reasoning capabilities of virtual assistants, particularly in situations where information must be retrieved from multiple sourcesSalveet al\.\([2024](https://arxiv.org/html/2607.23481#bib.bib5)\)\. In our context, where we need to handle multiple dialectal variants and types of datasets, adopting a multi\-agent approach was a natural choice\. Figure[2](https://arxiv.org/html/2607.23481#S4.F2)presents the overall pipeline that summarizes the core of our reasoning and information retrieval system\.

![Refer to caption](https://arxiv.org/html/2607.23481v1/x2.png)Figure 2:RAG Pipeline\.To address the challenges of handling multiple Comorian dialects and heterogeneous data sources, we designed a multi\-agent system that orchestrates specialized agents, each responsible for a specific type of query or knowledge domain\. This architecture, illustrated in Figure[2](https://arxiv.org/html/2607.23481#S4.F2), enables flexible and robust information retrieval by combining vector search, knowledge graphs and external web resources\.

### 4\.1Overall Pipeline

The system operates as follows: when a user submits a query, it is first processed by a central Retrieval Agent\. This agent acts as an orchestrator: it analyzes the query and determines which specialized agent should be invoked to provide the most accurate and comprehensive answer\. The response is then generated by an LLM based on the retrieved information and returned to the user\.

### 4\.2Core Components and Specialized Agents

The system comprises five main specialized agents, each leveraging different tools and knowledge bases:

- •Retrieval Agent: As the entry point for all queries, this agent is responsible for query understanding and dynamic routing\. It decides whether to forward the request to a single specialist or to combine results from multiple agents\. It relies on an LLM for intent classification and on a Vector Search engine to retrieve relevant passages from a precomputed embedding database covering all dialectal variants\.
- •Translator/Proverb Agent: This agent handles two closely related tasks\. For translation queries, it accesses bilingual dictionaries \(shiNgazidja\-English, shiMwali\-English\) and phrase collections\. For the queries related to proverb, we use a dedicated graph database that stores proverbs along with their meanings, cultural contexts and translations\. The graph structure allows for semantic navigation \(e\.g\., finding proverbs related to a specific theme\)\.
- •Teacher Agent: Focused on pedagogical needs, this agent provides grammar explanations, conjugation tables and language learning resources\. It draws upon the grammar manuals \(shiMaore and shiNdzuani\) and can generate structured lessons or exercises\. It also maintains a database of common learner errors and misconceptions\.
- •Culture Agent: For queries that fall outside the scope of other agents or require general information not covered by specialized knowledge bases, this agent performs live web searches\. It serves as a fallback mechanism, ensuring that the system can still provide relevant answers even when internal resources are insufficient\.

### 4\.3Technical Implementation

The system is built around several key technologies\. For embeddings and vector search, all textual resources including dictionaries, proverbs, phrase collections and grammar lessons, were embedded using Qwen3 EmbeddingZhanget al\.\([2025](https://arxiv.org/html/2607.23481#bib.bib4)\), a multilingual sentence transformer\. This model was selected for two main reasons\. First, it is among the top\-performing embedding models on recent benchmarks, offering high\-quality semantic representations across multiple languages\. Second, with only 2 billion parameters, it is lightweight enough to run in resource\-constrained environments, making it suitable for deployment in regions with limited computational infrastructure\. The model was deployed locally using Ollama, a lightweight framework for running large language models efficiently on consumer hardware\. The resulting embeddings were indexed in FAISS, a vector databaseDouzeet al\.\([2025](https://arxiv.org/html/2607.23481#bib.bib3)\)for efficient similarity search, enabling rapid retrieval of relevant passages based on query semantics rather than keyword matching alone\.

Regarding the knowledge graph, proverbs and cultural knowledge are stored in a graph database using Neo4j\. In this structure, nodes represent proverbs, definitions and everyday life entities, while edges capture semantic and thematic relationships\. This design allows for nuanced navigation and reasoning across the knowledge base, which is particularly valuable for answering queries\.

For orchestration and reasoning, agent coordination is implemented using LangChain\. This handles prompt engineering, tool invocation, dynamic query routing and response synthesis\. The core reasoning and language generation tasks are performed using Groq’s ultra\-fast inference platform, which hosts gpt\-oss, an open\-source LLM optimized for conversational AI\. Additionally, the Culture Agent can perform live searches using DuckDuckGo when information is unavailable internally\. Retrieved web content is summarized by gpt\-osss before integration into the final response\.

## 5Evaluation

### 5\.1Retrieval Performance

For this evaluation, we set a target of testing 500 queries, divided among five evaluators with 100 queries each\. Each evaluator categorized their queries according to the following four areas: Proverb Interpretation, Grammar Explanations, Vocabulary Lookup and Cultural Questions\. We measured the following metrics, often used to evaluate the quality of a recommendation systemJadon and Patil \([2024](https://arxiv.org/html/2607.23481#bib.bib2)\)and more recently for evaluating a RAG systemGanet al\.\([2025](https://arxiv.org/html/2607.23481#bib.bib1)\):

- •Precision@k: This metric measures the proportion of relevant documents among the topkresults\.
- •Recall@k: This metric measures the proportion of relevant items retrieved out of all existing relevant items\. In other words, if there arepprelevant items in total and among the topkkresults we findqqrelevant items, then Recall@k isq/pq/p\.
- •Mean Reciprocal Rank \(MRR\): The average of the reciprocal ranks of the first relevant result\. It is computed as follows: M​R​R=1N​∑i=1N1rankiMRR=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\frac\{1\}\{\\text\{rank\}\_\{i\}\}whereNNis the total number of queries andranki\\text\{rank\}\_\{i\}is the rank position of the first relevant result for queryii\.

Table[2](https://arxiv.org/html/2607.23481#S5.T2)presents retrieval performance across different query categories\. The system performs best on vocabulary lookup, which benefits from direct matches in dictionary entries\. Grammar explanations and proverb interpretation yield lower scores due to the complexity of these queries and the variability in how concepts are expressed across grammar manuals\. Cultural questions are the most challenging, as they often require synthesis from multiple sources or reliance on web search\.

Table 2:Retrieval performance by query category
### 5\.2Response Quality

The goal here is to evaluate the quality of the responses provided by the LLM\. To do this, we revisited the 500 queries used during the retrieval evaluation to examine the responses generated by the LLM\. These responses were evaluated using a 5\-point Likert scale along the following three axes:

- •Accuracy: Is the information factually correct?
- •Completeness: Does the response fully address the query?
- •Clarity: Is the response easy to understand for a learner?

Table[3](https://arxiv.org/html/2607.23481#S5.T3)summarizes the average scores\. The evaluators noted that vocabulary and grammar queries were generally well\-handled, while proverb interpretations sometimes lacked the depth of cultural context that a human expert would provide\. Cultural questions occasionally suffered from incomplete or overly general answers when relying on web search\.

Table 3:Human evaluation of response quality \(average scores, 1\-5 scale\)
### 5\.3Case Studies

While the quantitative metrics presented in the previous section provide a high\-level overview of system performance, they do not capture the nuances of individual interactions\. To complement these findings, we present four representative case studies \(See Figures[3](https://arxiv.org/html/2607.23481#S5.F3),[4](https://arxiv.org/html/2607.23481#S5.F4),[5](https://arxiv.org/html/2607.23481#S5.F5)and[6](https://arxiv.org/html/2607.23481#S5.F6)\) that illustrate the system’s behavior in practice\. These examples span the four different query categories analyzed previously\. Each case concludes with a brief discussion of strengths and limitations observed during the interaction\.

Figure 3:Vocabulary Query \(shiNgazidja\)User:Translate the word “hospital” in shiNgazidja\.

Assistant: Word:lopitali / lapitali \(malapitali\)Meaning:“hospital”Grammatical class:noun \(undefined\)Noun class:9‑10, 5‑6Origin:French \(fr\)

Analysis:The system correctly retrieves the dictionary entry with morphological variations \(singular/plural\) and etymology\.

Figure 4:Proverb InterpretationUser:I want a proverb about loss and its meaning\.

Assistant: Shikomori proverb:Namdodjipviwa ne shahanyu owakati mlonisho\.Meaning:We only realize the value of something after we have lost it; therefore, we must appreciate it while it is still there\.Topic:Appreciate the presentCategory:Time & Present

Analysis:The system successfully retrieves a relevant proverb based on semantic matching of “loss”\. The response includes both the original proverb and its French translation, along with thematic tags\. The informal query syntax \(“i want a proverb about loss”\) is correctly interpreted\.

Figure 5:Cultural QuestionsUser:What is the best way to greet an elderly person in ShiNdzuani?

Assistant:The most appropriate greeting for an older person in ShiNdzuani is:Mungu nahunusuru\.This is the customary salutation used when addressing someone older than yourself\.

Analysis:The system correctly identifies the culturally appropriate greeting for addressing elders in ShiNdzuani\. This demonstrates its ability to capture social norms and respectful language use, which is essential for learners navigating Comorian cultural contexts\. However, the response could be enriched with additional context, such as the literal meaning of the phrase \(“May God protect you”\) or the expected response from the elder\.

Figure 6:Grammar Explanation \(shiNdzuani\)User:How do noun classes work in shiNdzuani?

Assistant:In shiNdzuani, every noun belongs to anoun class\. The class is signalled by aprefixon the noun \(and on the plural form of the noun\)\. That same prefix is also used by adjectives, numbers, demonstratives and verbs so that theyagreewith the noun\.How the class is recognised:•Look at the singular prefix of the noun\.•The corresponding plural prefix tells you which class it is \(the “singular‑plural prefix correspondence”\)\.•Example: the singular prefixm\-\(ormw\-\) pairs with the plural prefixwa\-\(this is class 1‑2\)\.Examples:Agreement:•An adjective placed after the noun must take the same class prefix:gari kubwa\(big car\)\. If the noun were plural, the adjective would also change:magari makubwa\.•Numbers and ordinal numbers behave the same way; most change form to match the class, except a few invariable ones \(e\.g\.,muhimu,sita,saba,shendra,kumi\)\.•Verbs also carry the noun‑class prefix so that the subject and verb agree in class\.So, to use shiNdzuani correctly you:1\.Identify the noun’s class by its singular‑plural prefix pair\.2\.Apply the same prefix to any modifying adjective, number, demonstrative, or verb\.3\.If a word is listed without a prefix change \(e\.g\., the invariable adjectives above\), it does not change with the noun class\.

Analysis:This response demonstrates the system’s ability to provide structured grammatical explanations\. It correctly identifies noun class mechanisms and provides clear examples with agreement rules\.

### 5\.4Discussion of Limitations

Our evaluation reveals several limitations:

- •Coverage gaps: Despite our efforts, some dialectal variants and specialized vocabulary remain underrepresented, particularly for shiNdzuani and shiMwali\.
- •Cultural depth: Proverb interpretations and cultural explanations sometimes lack the nuanced understanding that a human expert would provide\.
- •Web search quality: When relying on web search, the system occasionally retrieves low\-quality or irrelevant information, affecting response accuracy\.
- •Evaluation scale: Our human evaluation involved only five annotators; a larger study with more diverse participants \(learners, teachers, elders\) would provide more robust insights\.

Despite these limitations, the results demonstrate thatMwandoprovides a solid foundation for computer\-assisted language learning for shiKomori, with particular strength in vocabulary and grammar support\.

## 6Conclusion and Future Work

In this paper, we introducedMwando, a virtual educational assistant for shiKomori, supporting its four dialectal variants\. We assembled a multi\-source corpus comprising phrases, proverbs, dictionaries and grammar lessons and developed a multi\-agent architecture combining vector search, a knowledge graph and web search fallback\. The system leverages lightweight embeddings \(Qwen3 via Ollama\) and fast reasoning \(Groq with gpt\-oss, an open\-source LLM\)\.

Evaluation on 500 queries showed strong performance on vocabulary lookup and grammar explanations, with qualitative case studies illustrating both capabilities and current limitations, particularly in cultural depth and proverb interpretation\.

Coverage remains uneven across dialects and web search quality is variable\. Future work will focus on expanding resources for shiNdzuani and shiMwali, enriching cultural knowledge with expert input, improving web search filtering and deployingMwandoin real\-world educational settings\. We hope this work contributes to the preservation of Comorian heritage and serves as a blueprint for AI support in other low\-resource languages\.

## References

- Y\. Abdillahi \(2012\)La diaspora de la Grande Comore à Marseille et son apport sur le développement de l’île\.Theses,Université de la Réunion\.External Links:[Link](https://theses.hal.science/tel-01206102)Cited by:[§2](https://arxiv.org/html/2607.23481#S2.p3.1)\.
- M\. Abdourahamane, C\. Boitet, V\. Bellynck, L\. Wang, and H\. Blanchon \(2016\)Construction d’un corpus parallèle français\-comorien en utilisant de la TA français\-swahili\.InTALAf \(Traitement Automatique des Langues africaines\),Paris, France\.External Links:[Link](https://hal.science/hal-01992871)Cited by:[§3\.1](https://arxiv.org/html/2607.23481#S3.SS1.p2.1)\.
- I\. Adebara, H\. O\. Toyin, N\. T\. Ghebremichael, A\. A\. Elmadany, and M\. Abdul\-Mageed \(2025\)Where are we? evaluating LLM performance on African languages\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 32704–32731\.External Links:[Link](https://aclanthology.org/2025.acl-long.1572/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1572),ISBN 979\-8\-89176\-251\-0Cited by:[§3\.1](https://arxiv.org/html/2607.23481#S3.SS1.p2.1)\.
- Agence Française de Développement \(2025\)Using AI to improve language learning in Senegal\.Agence Française de Développement\.Note:[https://www\.afd\.fr/en/AI\-for\-language\-learning\-in\-senegal](https://www.afd.fr/en/AI-for-language-learning-in-senegal)Accessed: 2026\-02\-14Cited by:[§2](https://arxiv.org/html/2607.23481#S2.p1.1)\.
- M\. Ahmed Chamanga \(2022\)ShiKomori, the bantu language of the comoros: status and perspectives\.InHandbook of Language Policy and Education in Countries of the Southern African Development Community \(SADC\),pp\. 79–98\.External Links:ISBN 9789004508057,[Link](http://dx.doi.org/10.1163/9789004516724_006),[Document](https://dx.doi.org/10.1163/9789004516724%5F006)Cited by:[§3\.1](https://arxiv.org/html/2607.23481#S3.SS1.p1.1)\.
- A\. Chauvet \(2015\)Statuts des langues et éducation de base aux comores\.Rev\. int\. d éduc\. Sèvres70\(70\),pp\. 77–84\.Cited by:[§2](https://arxiv.org/html/2607.23481#S2.p2.1),[§3\.1](https://arxiv.org/html/2607.23481#S3.SS1.p1.1)\.
- F\. Colace, R\. Gaeta, A\. Lorusso, M\. Pellegrino, and D\. Santaniello \(2025\)New ai challenges for cultural heritage protection: a general overview\.Journal of Cultural Heritage75,pp\. 168–193\.External Links:ISSN 1296\-2074,[Link](http://dx.doi.org/10.1016/j.culher.2025.07.019),[Document](https://dx.doi.org/10.1016/j.culher.2025.07.019)Cited by:[§1](https://arxiv.org/html/2607.23481#S1.p1.1)\.
- R\. S\. DANIEL \(2024\)Le shikomor pour enseignement / apprentissage du français langue étrangère : issue interculturelle de l’insularité de ndzuani\.Revue Internationale du Chercheur5\(2\)\.External Links:[Link](https://www.revuechercheur.com/index.php/home/article/view/984)Cited by:[§2](https://arxiv.org/html/2607.23481#S2.p2.1)\.
- M\. Douze, A\. Guzhva, C\. Deng, J\. Johnson, G\. Szilvasy, P\. Mazaré, M\. Lomeli, L\. Hosseini, and H\. Jégou \(2025\)The faiss library\.External Links:2401\.08281,[Link](https://arxiv.org/abs/2401.08281)Cited by:[§4\.3](https://arxiv.org/html/2607.23481#S4.SS3.p1.1)\.
- E\. Galaczi and R\. Luckin \(2024\)Generative AI and language education: opportunities, challenges and the need for critical perspectives\.Cambridge Papers in English Language EducationCambridge University Press & Assessment\.Note:Accessed: 2026\-02\-14External Links:[Link](https://www.cambridge.org/sites/default/files/media/documents/CPELE_Generative%20AI%20and%20Language%20Education%20Opportunities%20Challenges%20and%20the%20Need%20for%20Critical%20Perspectives_FINAL%20%281%29.pdf)Cited by:[§2](https://arxiv.org/html/2607.23481#S2.p1.1)\.
- A\. Gan, H\. Yu, K\. Zhang, Q\. Liu, W\. Yan, Z\. Huang, S\. Tong, and G\. Hu \(2025\)Retrieval augmented generation evaluation in the era of large language models: a comprehensive survey\.External Links:2504\.14891,[Link](https://arxiv.org/abs/2504.14891)Cited by:[§5\.1](https://arxiv.org/html/2607.23481#S5.SS1.p1.1)\.
- A\. Jadon and A\. Patil \(2024\)A comprehensive survey of evaluation techniques for recommendation systems\.InComputation of Artificial Intelligence and Machine Learning,pp\. 281–304\.External Links:ISBN 9783031714849,ISSN 1865\-0937,[Link](http://dx.doi.org/10.1007/978-3-031-71484-9_25),[Document](https://dx.doi.org/10.1007/978-3-031-71484-9%5F25)Cited by:[§5\.1](https://arxiv.org/html/2607.23481#S5.SS1.p1.1)\.
- V\. Koc \(2025\)Generative ai and large language models in language preservation: opportunities and challenges\.External Links:2501\.11496,[Link](https://arxiv.org/abs/2501.11496)Cited by:[§1](https://arxiv.org/html/2607.23481#S1.p1.1)\.
- M\. Lafon \(2007\)Le système Kamar\-Eddine : une tentative originale d’écriture du comorien en graphie arabe\.Ya Mkobe14\-15,pp\. 29–48\.External Links:[Link](https://shs.hal.science/halshs-00265704)Cited by:[§1](https://arxiv.org/html/2607.23481#S1.p3.1),[§3\.1](https://arxiv.org/html/2607.23481#S3.SS1.p2.1)\.
- Le Monde \(2024\)Comores : la diaspora installée en France dénonce son exclusion du scrutin présidentiel\.Le Monde\.Note:Accessed: 2026\-02\-14External Links:[Link](https://www.lemonde.fr/afrique/article/2024/01/10/comores-la-diaspora-installee-en-france-denonce-son-exclusion-du-scrutin-presidentiel_6210073_3212.html)Cited by:[§2](https://arxiv.org/html/2607.23481#S2.p3.1)\.
- T\. Liu, F\. Wang, and M\. Chen \(2024\)Rethinking tabular data understanding with large language models\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),K\. Duh, H\. Gomez, and S\. Bethard \(Eds\.\),Mexico City, Mexico,pp\. 450–482\.External Links:[Link](https://aclanthology.org/2024.naacl-long.26/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.26)Cited by:[§3\.3](https://arxiv.org/html/2607.23481#S3.SS3.p1.1)\.
- A\. M\. Naira, A\. Bahafid, Z\. Erraji, A\. Allak, M\. S\. Naoufal, and I\. Benelallam \(2025\)Preserving comorian linguistic heritage: bidirectional transliteration between the Latin alphabet and the Kamar\-eddine system\.InProceedings of the 9th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature \(LaTeCH\-CLfL 2025\),A\. Kazantseva, S\. Szpakowicz, S\. Degaetano\-Ortlieb, Y\. Bizzoni, and J\. Pagel \(Eds\.\),Albuquerque, New Mexico,pp\. 11–18\.External Links:[Link](https://aclanthology.org/2025.latechclfl-1.2/),[Document](https://dx.doi.org/10.18653/v1/2025.latechclfl-1.2),ISBN 979\-8\-89176\-241\-1Cited by:[§1](https://arxiv.org/html/2607.23481#S1.p2.1),[§1](https://arxiv.org/html/2607.23481#S1.p3.1),[§3\.1](https://arxiv.org/html/2607.23481#S3.SS1.p2.1)\.
- A\. M\. Naira, A\. Bahafid, Z\. Erraji, and I\. Benelallam \(2024\)Datasets creation and empirical evaluations of cross\-lingual learning on extremely low\-resource languages: a focus on comorian dialects\.InProceedings of the 18th Linguistic Annotation Workshop \(LAW\-XVIII\),S\. Henning and M\. Stede \(Eds\.\),St\. Julians, Malta,pp\. 140–149\.External Links:[Link](https://aclanthology.org/2024.law-1.14/)Cited by:[§1](https://arxiv.org/html/2607.23481#S1.p2.1),[§3\.1](https://arxiv.org/html/2607.23481#S3.SS1.p1.1),[§3\.1](https://arxiv.org/html/2607.23481#S3.SS1.p2.1)\.
- F\. Rotondo \(2016\)Cultural heritage as a key for the development of cultural and territorial integrated plans\.InCultural Territorial Systems,pp\. 21–27\.External Links:ISBN 9783319207537,ISSN 2194\-3168,[Link](http://dx.doi.org/10.1007/978-3-319-20753-7_4),[Document](https://dx.doi.org/10.1007/978-3-319-20753-7%5F4)Cited by:[§1](https://arxiv.org/html/2607.23481#S1.p1.1)\.
- I\. G\. Roukiyat \(2026\)The incorporation of shikomori to improve ict comprehension, access, and uptake by comorian communities\.Digital Policy Studies4\(1\),pp\. 57–83\.External Links:ISSN 2791\-3597,[Link](http://dx.doi.org/10.36615/cn8fx733),[Document](https://dx.doi.org/10.36615/cn8fx733)Cited by:[§2](https://arxiv.org/html/2607.23481#S2.p2.1),[§2](https://arxiv.org/html/2607.23481#S2.p3.1)\.
- A\. Salve, S\. Attar, M\. Deshmukh, S\. Shivpuje, and A\. M\. Utsab \(2024\)A collaborative multi\-agent approach to retrieval\-augmented generation across diverse data\.External Links:2412\.05838,[Link](https://arxiv.org/abs/2412.05838)Cited by:[§4](https://arxiv.org/html/2607.23481#S4.p1.1)\.
- M\. Serva and M\. Pasquini \(2021\)The sabaki languages of comoros\.INDIAN OCEAN REVIEW OF SCIENCE AND TECHNOLOGY\.External Links:[Link](http://www.iorst.net/index.php/paper/view/10)Cited by:[§3\.1](https://arxiv.org/html/2607.23481#S3.SS1.p1.1)\.
- L\. Team, A\. Modi, A\. S\. Veerubhotla, A\. Rysbek, A\. Huber, A\. Anand, A\. Bhoopchand, B\. Wiltshire, D\. Gillick, D\. Kasenberg, E\. Sgouritsa, G\. Elidan, H\. Liu, H\. Winnemoeller, I\. Jurenka, J\. Cohan, J\. She, J\. Wilkowski, K\. Alarakyia, K\. R\. McKee, K\. Singh, L\. Wang, M\. Kunesch, M\. Pîslar, N\. Efron, P\. Mahmoudieh, P\. Kamienny, S\. Wiltberger, S\. Mohamed, S\. Agarwal, S\. M\. Phal, S\. J\. Lee, T\. Strinopoulos, W\. Ko, Y\. Gold\-Zamir, Y\. Haramaty, and Y\. Assael \(2025\)Evaluating gemini in an arena for learning\.External Links:2505\.24477,[Link](https://arxiv.org/abs/2505.24477)Cited by:[§2](https://arxiv.org/html/2607.23481#S2.p1.1)\.
- United Nations Development Programme and Ministry of Enterprises and Made in Italy \(2024\)Scaling language data ecosystems to drive industrial development growth\.External Links:[Link](https://cdn.prod.website-files.com/66e31d90ea60e260f5ea025f/68546ed270a71196702c0081_Community%20Paper_Final%20for%20AI%20Hub%20Launch%20-%20REVISED%20VERSION%20(1).pdf)Cited by:[§1](https://arxiv.org/html/2607.23481#S1.p1.1)\.
- H\. Wei, Y\. Sun, and Y\. Li \(2025\)DeepSeek\-ocr: contexts optical compression\.arXiv preprint arXiv:2510\.18234\.Cited by:[2nd item](https://arxiv.org/html/2607.23481#S3.I1.i2.p1.1)\.
- S\. Wild \(2025\)AI models are neglecting african languages — scientists want to change that\.Nature\.External Links:ISSN 1476\-4687,[Link](http://dx.doi.org/10.1038/d41586-025-02292-5),[Document](https://dx.doi.org/10.1038/d41586-025-02292-5)Cited by:[§1](https://arxiv.org/html/2607.23481#S1.p1.1)\.
- M\. Zaim, S\. Arsyad, B\. Waluyo, H\. Ardi, Muhd\. Al Hafizh, M\. Zakiyah, W\. Syafitri, A\. Nusi, and M\. Hardiah \(2025\)Generative ai as a cognitive co\-pilot in english language learning in higher education\.Education Sciences15\(6\),pp\. 686\.External Links:ISSN 2227\-7102,[Link](http://dx.doi.org/10.3390/educsci15060686),[Document](https://dx.doi.org/10.3390/educsci15060686)Cited by:[§2](https://arxiv.org/html/2607.23481#S2.p1.1)\.
- Y\. Zhang, M\. Li, D\. Long, X\. Zhang, H\. Lin, B\. Yang, P\. Xie, A\. Yang, D\. Liu, J\. Lin, F\. Huang, and J\. Zhou \(2025\)Qwen3 embedding: advancing text embedding and reranking through foundation models\.External Links:2506\.05176,[Link](https://arxiv.org/abs/2506.05176)Cited by:[§4\.3](https://arxiv.org/html/2607.23481#S4.SS3.p1.1)\.

Similar Articles

WebMCP: Teaching Your Website to Talk to AI Agents

Hacker News Top

WebMCP is a proposed web standard developed by Google and Microsoft that allows websites to declare structured tools for AI agents to call directly, replacing fragile screen-scraping with stable interfaces.

@hwchase17: https://x.com/hwchase17/status/2071963622298050997

X AI KOLs Timeline

The article discusses the emerging pattern of 'wiki memory' for AI agents, where raw source data is intelligently compressed into a persistent, structured knowledge layer that agents can use efficiently. It compares this to basic RAG and gives examples like DeepWiki and LLM Wiki.