TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

arXiv cs.CL Papers

Summary

The paper introduces TranslatePsy-AfriSLM, an open-source collection of machine translation resources for 19 Sub-Saharan African languages, demonstrating that fine-tuned small language models with filtered synthetic data outperform much larger models like TranslateGemma-27B and Qwen3.5-122B-A10B.

arXiv:2608.18655v1 Announce Type: new Abstract: The rapid progress in Artificial Intelligence has largely bypassed African languages, creating a digital divide that limits AI adoption on the continent. Recent open-source LLMs systematically underperform on African machine translation, while the lack of large-scale, high-quality, open-source parallel data has constrained the development of competitive small language models (SLMs). We introduce *TranslatePsy-AfriSLM*, a collection of open-source MT resources for 19 Sub-Saharan African languages, including curated parallel data, African-specialized synthetic data, and a family of fine-tuned SLMs. Our empirical study shows that unified quality-estimation filtering removes up to 96% of training tokens without degrading quality, and that filtered synthetic data dominates the quality-efficiency Pareto frontier. Fine-tuned on the resulting data mixture, TranslatePsy-AfriSLM outperforms substantially larger systems, including TranslateGemma-27B and Qwen3.5-122B-A10B, with as few as 0.8B parameters.
Original Article
View Cached Full Text

Cached at: 08/20/26, 10:15 AM

# High-Quality Data Scaling For Low-Resource Machine Translation
Source: [https://arxiv.org/html/2608.18655](https://arxiv.org/html/2608.18655)
Patrik Lambert11footnotemark:1Jihye Back11footnotemark:1Amril NazirAffiliation:Tether AI ResearchAffiliation:\{milan\.gritta, patrik\.lambert, jihye\.back, amril\.nazir\}@tether\.io

###### Abstract

The rapid progress in Artificial Intelligence has largely bypassed African languages, creating a digital divide that limits AI adoption on the continent\. Recent open\-source LLMs systematically underperform on African machine translation, while the lack of large\-scale, high\-quality, open\-source parallel data has constrained the development of competitive small language models \(SLMs\)\. We introduceTranslatePsy\-AfriSLM, a collection of open\-source MT resources for 19 Sub\-Saharan African languages, including curated parallel data, African\-specialized synthetic data, and a family of fine\-tuned SLMs\. Our empirical study shows that unified quality\-estimation filtering removes up to 96% of training tokens without degrading quality, and that filtered synthetic data dominates the quality\-efficiency Pareto frontier\. Fine\-tuned on the resulting data mixture, TranslatePsy\-AfriSLM outperforms substantially larger systems, including TranslateGemma\-27B and Qwen3\.5\-122B\-A10B, with as few as 0\.8B parameters\.

## 1Introduction

The AI underinvestment on the African continent[30](https://arxiv.org/html/2608.18655#bib.bib15);[12](https://arxiv.org/html/2608.18655#bib.bib39);[18](https://arxiv.org/html/2608.18655#bib.bib40)has created a significant adoption barrier for over a billion people, preventing them from fully exploiting the productivity and collaboration benefits that AI can offer[25](https://arxiv.org/html/2608.18655#bib.bib38)\. A prime example of this is a lack of performant SLMs for Machine Translation \(MT\), a critical utility for facilitating cross\-border communication, trade, and education[42](https://arxiv.org/html/2608.18655#bib.bib41);[27](https://arxiv.org/html/2608.18655#bib.bib42)\. However, most frontier open\-source LLMs such as Apertus[17](https://arxiv.org/html/2608.18655#bib.bib44), Qwen3[50](https://arxiv.org/html/2608.18655#bib.bib3), TranslateGemma[15](https://arxiv.org/html/2608.18655#bib.bib2), Hunyuan\-MT[53](https://arxiv.org/html/2608.18655#bib.bib20)or Qwen3\.5[45](https://arxiv.org/html/2608.18655#bib.bib19)systematically underperform on African machine translation \(see Figure[1](https://arxiv.org/html/2608.18655#S1.F1)\) while incurring substantial running costs due to large parameter counts\.

![Refer to caption](https://arxiv.org/html/2608.18655v1/figures/size_vs_performance_bouquet.png)Figure 1:TranslatePsy\-AfriSLM surpasses frontier LLMs\(Qwen3\.5\-122B\-A10B, TranslateGemma\-27B\)\. SSA\-COMET scores shown on BOUQuET benchmark\.![Refer to caption](https://arxiv.org/html/2608.18655v1/figures/data_pipeline.png)Figure 2:TranslatePsy\-AfriSLM data pipeline\.Parallel and monolingual sources are unified through preprocessing, QE filtering, calibrated on the Human Mix, finally the optimal Open\-Source and Synthetic mixes are selected\.Furthermore, efficiently adapting existing SLMs to support this task remains challenging due to the scarcity of high\-quality, open\-source data in sufficient quantities, see Background \(§[2](https://arxiv.org/html/2608.18655#S2)\) for a detailed survey\. Large internet repositories are highly unstructured and noisy, the training signal is sparse and the token budget dramatically inflated\. Existing curated datasets and human\-quality translations are too small, offering only modest improvements\. Therefore, to support efficient model adaptation for low\-resource machine translation, we introduceTranslatePsy\-AfriSLM111[https://huggingface\.co/collections/qvac/translatepsy\-afrislm](https://huggingface.co/collections/qvac/translatepsy-afrislm), an open\-source African MT resource comprising high\-quality parallel data for 19 Sub\-Saharan African languages and a family of highly\-capable SLMs, shown in Figure[1](https://arxiv.org/html/2608.18655#S1.F1)\. Our approach is systematically validated on key benchmarks, with deep insights and analyses that can guide future research\. TranslatePsy\-AfriSLM exceeds African translation skills of frontier LLMs such as TranslateGemma\-27B and Qwen3\.5\-27B \(and 122B\-A10B\) with only 0\.8B parameters\. Our smallest model even outperforms the dedicated NLLB encoder\-decoder models while preserving conversational capabilities \(demo in Figure[13](https://arxiv.org/html/2608.18655#A4.F13)\)\.

## 2Background

### 2\.1African MT Training Resources

Existing African MT datasets reveal a persistent trade\-off between scale, coverage, and translation quality\. We broadly group them into 3 categories: large but unstructured repositories, medium\-sized curated corpora, and small, human\-quality datasets\.

#### Large but Unstructured\.

Data repositories such as OPUS[47](https://arxiv.org/html/2608.18655#bib.bib14), MALA[19](https://arxiv.org/html/2608.18655#bib.bib18), WMT22[3](https://arxiv.org/html/2608.18655#bib.bib17)and Fine Translations[35](https://arxiv.org/html/2608.18655#bib.bib16)contain large quantities of parallel data, however, the sentence pairs occur in arbitrary sizes, number of languages, translation directions, suffer from high duplication rates, test data contamination and noisy texts\. NLLB[44](https://arxiv.org/html/2608.18655#bib.bib21)leveraged approx\. 18B sentence pairs across 200 languages, but did not release a fixed, high\-quality corpus suitable for token\-efficient adaptation of African MT models\.

#### Medium\-Sized and Curated\.

AfriNLLB222[https://huggingface\.co/datasets/AfriNLP/AfriNLLB\-train](https://huggingface.co/datasets/AfriNLP/AfriNLLB-train)is currently the only readily available, medium\-sized, curated dataset for machine translation[26](https://arxiv.org/html/2608.18655#bib.bib4)\. However, it covers only 9 African languages, and approximately 50% of its∼\\sim3M pairs include Arabic or European languages\. AfriqueLLM[51](https://arxiv.org/html/2608.18655#bib.bib5)adapts frontier LLMs for African languages through continued pretraining\. While its general\-purpose scope significantly improves MT, its training data was not publicly released \(a gap we aim to fill\)\. We report comparisons with AfriqueLLM models in Figure[1](https://arxiv.org/html/2608.18655#S1.F1)and SLMs trained on AfriNLLB data in Section[5\.4](https://arxiv.org/html/2608.18655#S5.SS4)\.

#### Small but Human\-Quality\.

Several datasets are available in this subset, e\.g\. MMT\-Africa[14](https://arxiv.org/html/2608.18655#bib.bib45), AfriDOC\-MT[7](https://arxiv.org/html/2608.18655#bib.bib12), SMOL[9](https://arxiv.org/html/2608.18655#bib.bib11), LAFAND\-MT[2](https://arxiv.org/html/2608.18655#bib.bib46), WMT24pp[10](https://arxiv.org/html/2608.18655#bib.bib10), MENYO\-MT[6](https://arxiv.org/html/2608.18655#bib.bib53), however, they occur in limited quantities and have uneven language coverage\. We show that despite their high\-quality translations, such mixtures yield only modest improvements \(§[5\.3](https://arxiv.org/html/2608.18655#S5.SS3)\), smaller than AfriNLLB and substantially smaller than our best TranslatePsy\-AfriSLM mixes\. This further motivates our large\-scale, quality\-focused data curation\.

### 2\.2Quality Estimation for Data Filtering

Afri\-COMET[49](https://arxiv.org/html/2608.18655#bib.bib6)and more recently, SSA\-COMET[24](https://arxiv.org/html/2608.18655#bib.bib35)were introduced as African\-centric, reference\-free alternatives to COMET[38](https://arxiv.org/html/2608.18655#bib.bib47), COMET\-KIWI[39](https://arxiv.org/html/2608.18655#bib.bib37)and MetricX[22](https://arxiv.org/html/2608.18655#bib.bib9)\. Several prior works have used them to filter parallel text in African languages[51](https://arxiv.org/html/2608.18655#bib.bib5);[48](https://arxiv.org/html/2608.18655#bib.bib49);[26](https://arxiv.org/html/2608.18655#bib.bib4)\. However, two questions remain unanswered: \(1\) the comparative usefulness of each metric*as a training data filter*, as opposed to an evaluation metric, validated on large\-scale African MT; and \(2\) whether a combination of individual metrics would perform more consistently, and if so, how to effectively combine multiple QE metrics \(§[3](https://arxiv.org/html/2608.18655#S3)\)\.

## 3TranslatePsy\-AfriSLM

We address the heterogeneity of African MT data with a multi\-step curation pipeline, shown in Figure[2](https://arxiv.org/html/2608.18655#S1.F2), applied to open\-source and synthetic sentence pairs \(§[3\.1](https://arxiv.org/html/2608.18655#S3.SS1)\)\. After standard structural preprocessing, we score raw pairs using Unified Quality Estimation \(§[3\.2](https://arxiv.org/html/2608.18655#S3.SS2)\), a robust z\-score that aggregates multiple QE metrics\. We then study data quantity selection methods for post\-training \(§[3\.3](https://arxiv.org/html/2608.18655#S3.SS3)\)\.

### 3\.1Data Sources

#### Parallel Data\.

We source data from four large, open\-source repositories: WMT22[3](https://arxiv.org/html/2608.18655#bib.bib17), MALA[19](https://arxiv.org/html/2608.18655#bib.bib18), OPUS[47](https://arxiv.org/html/2608.18655#bib.bib14)and Fine Translations[35](https://arxiv.org/html/2608.18655#bib.bib16)\. Unlike synthetic data, where we can control generation volume and direction, open\-source corpora have fixed and often highly uneven distributions across languages and translation directions\. As a result, these variables canvary significantlybetween sources\. Therefore, we cap each language pair to a maximum of 5M sentence pairs to mitigate ’winner\-takes\-all’ effects\. This process gives us approximately 427 million raw sentence pairs, see Table[4](https://arxiv.org/html/2608.18655#A1.T4)for a per\-language breakdown\.

#### Monolingual Data\.

We used the MADLAD\-400 corpus[23](https://arxiv.org/html/2608.18655#bib.bib1)as a source of data for synthetic data generation via machine translation, see "Synthetic Data Generation" below\. For each of the 19 African languages considered in this work, we processed all data available in the corpus\. For English, due to the extremely large amounts of data available, we randomly sampled 3\.6 million documents\. We report detailed, per\-language dataset statistics in Table[5](https://arxiv.org/html/2608.18655#A1.T5)\.

#### Preprocessing\.

This includes standard cleaning, language checks, deduplication, and test set decontamination\. Details are provided in Appendix[A\.1](https://arxiv.org/html/2608.18655#A1.SS1)\.

#### Synthetic Data Generation

To generate synthetic data, we translated the monolingual data with the NLLB\-3\.3B encoder\-decoder, selected as the teacher model based on our benchmarks in Table[17](https://arxiv.org/html/2608.18655#A4.T17)\. See Appendix[A\.1](https://arxiv.org/html/2608.18655#A1.SS1)for generation settings\.

### 3\.2Unified Quality Estimation

We aim to filter the open\-source and synthetic data with AfriCOMET,333[https://huggingface\.co/Masakhane/africomet\-qe\-stl\-1\.1](https://huggingface.co/Masakhane/africomet-qe-stl-1.1)SSA\-COMET444[https://huggingface\.co/McGill\-NLP/ssa\-comet\-mtl](https://huggingface.co/McGill-NLP/ssa-comet-mtl)\- according to[24](https://arxiv.org/html/2608.18655#bib.bib35), this model achieves better results thanssa\-comet\-qe\.and MetricX555[https://huggingface\.co/google/metricx\-24\-hybrid\-xl\-v2p6withoutreference\.](https://huggingface.co/google/metricx-24-hybrid-xl-v2p6withoutreference.)QE metrics666Note:These metrics are used for both quality estimationandreference\-based evaluation\. We use the names interchangeably depending on local context\.\. As we show later \(§[5\.1](https://arxiv.org/html/2608.18655#S5.SS1)\), no single quality estimator consistently performs best across all evaluation metrics\. This motivates aggregating all three estimators into a single score for more consistent filtering\. However, a naïve aggregation seems impractical, as the scores have distinct polarities \(higher\- versus lower\-is\-better\) and their ranges vary across QE models, language directions and training corpora\. To address this, we unify the quality estimates from AfriCOMET, SSA\-COMET and MetricX by mapping them into a sharedrobust z\-score\(Eq\.[2](https://arxiv.org/html/2608.18655#S3.E2)\), calibrated on the following dataset\.

#### Human Mix\.

We compiled ~352K high\-quality, human\-translated pairs from two datasets \(see Appendix[A\.2](https://arxiv.org/html/2608.18655#A1.SS2)for details\) for two distinct purposes: a\) to provide a human\-quality reference for the robust z\-score parameters, and b\) to show that high\-quality but limited\-scale human\-translated data alone is insufficient for effective post\-training adaptation\.

#### Average Robust z\-score\.

For each translation directionddand for each metricmm, we compute Human Mix calibration statistics: the medianx~d,m\\mathrm\{\\tilde\{x\}\_\{d,m\}\}and the Median Absolute Deviation \(MAD\):

MADd,m=median⁡\(\|xi−x~d,m\|\)\.\\mathrm\{\\operatorname\{MAD\}\_\{d,m\}=\\operatorname\{median\}\\\!\\left\(\\left\|x\_\{i\}\-\\tilde\{x\}\_\{d,m\}\\right\|\\right\)\.\}\(1\)Each candidate QE scorexi,d,m\\mathrm\{x\_\{i,d,m\}\}is then normalized against these statistics:

zi,d,m=sm⋅0\.6745​\(xi,d,m−x~d,m\)MADd,m\\mathrm\{z\_\{i,d,m\}=s\_\{m\}\\cdot\\frac\{0\.6745\\,\(x\_\{i,d,m\}\-\\tilde\{x\}\_\{d,m\}\)\}\{\\operatorname\{MAD\}\_\{d,m\}\}\}\(2\)wheresm∈\{\+1,−1\}\\mathrm\{s\_\{m\}\\in\\\{\+1,\-1\\\}\}flips lower\-is\-better metrics such as MetricX, so that higher values always indicate higher quality\. By calibrating each score against the same Human Mix statistics \(Eq\.[2](https://arxiv.org/html/2608.18655#S3.E2)\), this normalization places heterogeneous data sources, translation directions, and QE metrics on a common quality scale relative to human\-translated data\. They can now be aggregated by computing the average of the per\-metric normalized scores in Equation[3](https://arxiv.org/html/2608.18655#S3.E3)\. We adopt thisaverage robust z\-scoreas a unified QE metric, henceforth denoted simply asz¯\\mathrm\{\\bar\{z\}\}:

z¯i,d=13​∑m=13zi,d,m\\mathrm\{\\bar\{z\}\_\{i,d\}=\\frac\{1\}\{3\}\\sum\_\{m=1\}^\{3\}z\_\{i,d,m\}\}\(3\)This unified score provides a more consistent basis for filtering heterogeneous parallel data\.

#### Directional Filtering\.

Quality estimation scores are not symmetric with respect to theirargument order; consequently, for any training pair\(X,Y\)\\mathrm\{\(X,Y\)\}the scorez¯​\(X,Y\)\\mathrm\{\\bar\{z\}\(X,Y\)\}typically differs fromz¯​\(Y,X\)\\mathrm\{\\bar\{z\}\(Y,X\)\}\. The choice of scoring direction therefore determines which sentence pairs pass the filter\. We investigate the following strategies for filtering:

- •Aligned:z¯​\(X,Y\)\\mathrm\{\\bar\{z\}\(X,Y\)\}, trainX→Y\\mathrm\{X\\rightarrow Y\}
- •Reversed:z¯​\(Y,X\)\\mathrm\{\\bar\{z\}\(Y,X\)\}, trainX→Y\\mathrm\{X\\rightarrow Y\}
- •Mean:12​\(z¯​\(X,Y\)\+z¯​\(Y,X\)\)\\mathrm\{\\tfrac\{1\}\{2\}\\bigl\(\\bar\{z\}\(X,Y\)\+\\bar\{z\}\(Y,X\)\\bigr\)\}, trainX→Y\\mathrm\{X\\rightarrow Y\}

Thealignedstrategy scores each pair in the training direction, whereas thereversedstrategy scores it in the opposite direction\. Themeanstrategy averages the z\-scores from both directions, potentially providing a balanced alternative that we also explore\.

### 3\.3Data Quantity Selection

Once the QE strategy is determined, we need to investigate methods of selecting training data quantities: 1\)Thresholdfilters out sentence pairs below a givenz¯\\bar\{z\}score, 2\)TopNextends the threshold filter by limiting any single language pair to a maximum of top\-N examples to mitigate large imbalances between languages, 3\)Bidirectional expansionis an orthogonal data augmentation step that reverses each sentence pair\(X,Y\)→\(Y,X\)\\mathrm\{\\mathrm\{\(X,Y\)\\rightarrow\(Y,X\)\}\}to balance translation directions during training\.

### 3\.4Auxiliary Data Mixes

In addition to the open\-source and synthetic mixes, we include two auxiliary mixtures to maintain capabilities related to broader machine translation\.

#### Instruct Mix

is our multilingual instruction\-following dataset with approximately 50% African\-language content, totalling 4\.6M examples\. We include it to preserve general conversational abilities and to extend translation to multi\-turn settings\. The dataset composition is provided in Appendix[A\.3](https://arxiv.org/html/2608.18655#A1.SS3)\.

#### Asia\-Europe Mix

covers 38 languages \(∼\\sim24M examples\)\. We include it to preserve translation quality on medium\- and high\-resource languages and to study whether such data mitigates catastrophic forgetting, see §[A\.4](https://arxiv.org/html/2608.18655#A1.SS4)for dataset details\.

## 4Experimental Setup

### 4\.1Supervised Fine\-Tuning \(SFT\)

We post\-train Qwen3\.5[45](https://arxiv.org/html/2608.18655#bib.bib19)models using Supervised Fine\-Tuning \(SFT\), as they consistently outperform other general\-purpose SLMs on in our preliminary benchmarks\. Each model undergoes \(full\-parameter\) fine\-tuning for one epoch on the final training mixture selected from our data analyses in §[5\.2](https://arxiv.org/html/2608.18655#S5.SS2), with loss computed only on assistant tokens\. Details are provided in Appendix[B\.2](https://arxiv.org/html/2608.18655#A2.SS2)\.

![Refer to caption](https://arxiv.org/html/2608.18655v1/figures/summary_avg_p10_p50_p90_zscore.png)Figure 3:Comparison of data quality across data sources\.Each column shows the P10, P50, and P90 percentiles for AfriCOMET, SSA\-COMET, MetricX, which are normalised into a mean robust z\-score \(bottom row\)\.
### 4\.2Evaluation

#### Language groups\.

We evaluate MT on 19 languages covered in the TranslatePsy\-AfriSLM data mixes, labelledAfrica\-IID777Africa\-IID: Afrikaans \(afr\), Amharic \(amh\), Hausa \(hau\), Igbo \(ibo\), Kinyarwanda \(kin\), Lingala \(lin\), Luganda \(lug\), Malagasy \(plt/mlg\), Nyanja \(nya\), Oromo \(orm\), Nigerian Pidgin \(pcm\), Shona \(sna\), Somali \(som\), Southern Sotho \(sot\), Swahili \(swh/swa\), Tswana \(tsn\), Wolof \(wol\), Xhosa \(xho\), Yoruba \(yor\), Zulu \(zul\); see[https://en\.wikipedia\.org/wiki/ISO\_639\-3](https://en.wikipedia.org/wiki/ISO_639-3)\., and an additional set of 8 unseen languages, labelledAfrica\-OOD888Africa\-OOD: Nigerian Pidgin \(pcm\), Sudanese Arabic \(apd\), Akan \(aka\), Tamazight \(ber\), Kituba \(ktu\), Bambara \(bam\), Sepedi \(nso\) and Mooré \(mos\)\., to assess generalisation to out\-of\-domain languages\.

#### Benchmarks\.

We evaluate on three translation benchmarks: Flores\-200[44](https://arxiv.org/html/2608.18655#bib.bib21), BOUQuET[46](https://arxiv.org/html/2608.18655#bib.bib28), and Smol[9](https://arxiv.org/html/2608.18655#bib.bib11)\. Flores\-200 is the most widely used benchmark \(devtestsplit, 1,012 sentences\), enabling comparison with prior work\. BOUQuET is the most recent, linguist\-curated multi\-way benchmark \(testsplit, 854 sentences\) with broad typological coverage\. Smol \(smolsentsplit, 863 sentences\) provides professionally translated data across 89 languages, used exclusively for evaluation\. Each benchmark covers all 19 Africa\-IID languages\.

#### Metrics\.

We adopt four complementary metrics:COMET\-22999[https://huggingface\.co/Unbabel/wmt22\-comet\-da](https://huggingface.co/Unbabel/wmt22-comet-da)[37](https://arxiv.org/html/2608.18655#bib.bib7), a neural metric trained on human direct assessments, abbreviated toC22;SSA\-COMET101010[https://huggingface\.co/McGill\-NLP/ssa\-comet\-mtl](https://huggingface.co/McGill-NLP/ssa-comet-mtl)[24](https://arxiv.org/html/2608.18655#bib.bib35), an AfroXLMR\-based COMET variant targeting African languages, abbreviated toSSA;MetricX111111[https://huggingface\.co/google/metricx\-24\-hybrid\-xl\-v2p6](https://huggingface.co/google/metricx-24-hybrid-xl-v2p6)[21](https://arxiv.org/html/2608.18655#bib.bib8), an mT5\-based regression metric fine\-tuned on DA and MQM ratings, abbreviated toMX; andChrF\+\+121212Sacrebleu python library withchar\_order=6,word\_order=2,beta=2,lowercase=False,whitespace=False,eps\_smoothing=False[36](https://arxiv.org/html/2608.18655#bib.bib43), a character and word\-based n\-gram F\-score robust to morphological variation\. COMET\-22 and SSA\-COMET are reported on a 0–1 scale \(higher is better\), ChrF\+\+ on a 0–100 scale \(higher is better\), and MetricX on a 0–25 scale \(lower is better\)\.

#### Baselines

We benchmark our SLMs against\(1\) General\-purpose LLMs:Qwen3[50](https://arxiv.org/html/2608.18655#bib.bib3)and Qwen3\.5[45](https://arxiv.org/html/2608.18655#bib.bib19), two strong general\-purpose multilingual LLM families, and Apertus[17](https://arxiv.org/html/2608.18655#bib.bib44), a massively multilingual model covering 1,800\+ languages\. Among these, Qwen3\.5 achieves the strongest performance, motivating its use as our backbone and enabling a direct measurement of gains from our methodology\.\(2\) Dedicated translation models:NLLB[44](https://arxiv.org/html/2608.18655#bib.bib21), an encoder–decoder model trained on large\-scale parallel data; AfriNLLB[26](https://arxiv.org/html/2608.18655#bib.bib4), an African\-centric variant of NLLB; TranslateGemma[15](https://arxiv.org/html/2608.18655#bib.bib2)and Hunyuan\-MT[53](https://arxiv.org/html/2608.18655#bib.bib20), decoder\-only LLMs post\-trained for translation\.\(3\) African language\-specialized models:AfriqueLLM[51](https://arxiv.org/html/2608.18655#bib.bib5), which adapts Gemma\-3, Qwen3, and LLaMA\-3\.1 via continued pretraining on approximately 26B tokens\.

## 5Results and Analysis

### 5\.1Quality Estimation: Which is Best?

We first study which quality estimation strategy is most effective for filtering training data\. All experiments in this section use English\-to\-African synthetic data, the only corpus large enough to support controlled comparisons across all 19 language pairs, with over 135 million examples\.

#### No Single Estimator Is Optimal\.

We compare SSA\-COMET, AfriCOMET, and MetricX as individual quality estimators for filtering training data\. For each estimator, we score the full training pool in both translation directions, select the top 2 million examples \(per language\), and use them to fine\-tune Qwen3\.5\-2B\. Results on BOUQuET test set are shown in Table[1](https://arxiv.org/html/2608.18655#S5.T1)\. Flores\-200 and Smol follow the same pattern, as shown in Tables[6](https://arxiv.org/html/2608.18655#A1.T6)and[7](https://arxiv.org/html/2608.18655#A1.T7)\. We observe that each estimator performs best when evaluated by its corresponding evaluation metric\. For instance, SSA\-COMET estimator consistently achieves the top SSA\-COMET score, and the same holds for MetricX\. Similarly, the AfriCOMET estimator outperforms MetricX QE when evaluated via COMET variants\. Consequently, no single estimator dominates across all metrics\. However, thez¯\\bar\{z\}score does provide a balanced alternative, achievingat least the second best score across all metrics\. We therefore adopt it as our QE metric\.

Quality estimatorC22SSAMXChrF\+\+eng\-xx \(English→\\rightarrowAfrican language\(s\)\)z\-score0\.7650\.6433\.8050\.1SSA\-COMET0\.7660\.6473\.9650\.5AfriCOMET0\.7630\.6374\.0349\.8MetricX0\.7600\.6323\.6249\.6xx\-eng \(African language\(s\)→\\rightarrowEnglish\)z\-score0\.7860\.6064\.5952\.7SSA\-COMET0\.7850\.6084\.8253\.2AfriCOMET0\.7830\.6024\.7152\.4MetricX0\.7820\.5984\.5052\.1Table 1:Translation quality with individual versus unified QE metricsover BOUQuET test set\.
#### Aligned QE Is Best\.

We evaluate whether QE filtering should be applied in the training direction \(aligned\), in the opposite of training direction \(reversed\) or both \(mean\)\. This issue arises in back\-translation workflows, where synthetic pairs may be generated and filtered in one direction, then reversed to form training examples in the opposite direction\. In such cases, the QE scoring direction becomesreversedrelative to the final training direction\. To test this, we compare otherwise identical models filtered using thealigned,reversed, andmeanstrategies defined in §[3\.2](https://arxiv.org/html/2608.18655#S3.SS2)\. Results in Table[2](https://arxiv.org/html/2608.18655#S5.T2), averaged across translation directions and test sets, show that reversed filtering substantially degrades performance, especially on MetricX \(−12\.0%\-12\.0\\%\) and SSA\-COMET \(−3\.1%\-3\.1\\%\), while the mean strategy is closer but still generally below the aligned reference\. This indicates that QE filteringshould be aligned with the final training direction, even when synthetic data is generated through back\-translation\. We thus use aligned QE filtering for the remainder of the paper\.

![Refer to caption](https://arxiv.org/html/2608.18655v1/figures/pareto_frontier_ssa_vs_tokens__bouquet.png)Figure 4:The Pareto frontier of performance \(SSA\-COMET\) versus training tokens on BOUQuET\.Each shape is a fine\-tuned model configuration\. Colours indicate data source, shapes indicate the filtering strategy, and the shape fill indicates the threshold \(the fuller the shape, the higher the threshold\)\. Detailed plots for each metric and each dataset are shown in Figure[12](https://arxiv.org/html/2608.18655#A3.F12)\. The full training budgets and threshold details are reported in Table[15](https://arxiv.org/html/2608.18655#A4.T15)\.DirectionC22SSAMXChrF\+\+Aligned \(Ref\.\)0000Mean\-0\.4%\-0\.9%\-3\.9%\+0\.7%Reversed\-1\.5%\-3\.1%\-12\.0%\+0\.1%

Table 2:Training versus QE scoring directions\.Percentages show average differences toalignedreference\.

### 5\.2How to Select Data?

With unified QE established, we study how training performance changes as a function of data quality and quantity\. We sweep the z\-score thresholds and combine them with the three strategies introduced in §[3\.3](https://arxiv.org/html/2608.18655#S3.SS3)across open\-source, synthetic and combined data sources\. For each mix, we fine\-tune Qwen3\.5\-2B under the same recipe\. Our key observations:

1. 1\.Quality filtering is highly effective\.Figure[4](https://arxiv.org/html/2608.18655#S5.F4)\(bottom right\) shows the “Open\-source \(Unfiltered\)” model trained only on raw sentence pairs, totalling 44\.93B tokens\. Its lower relative quality can clearly be seen in Figure[3](https://arxiv.org/html/2608.18655#S4.F3)\. In contrast, a filtered configuration \(bottom left\) reaches a comparable SSA\-COMET score \(0\.530 vs\. 0\.528\) with only 1\.76B tokens—a 96% reduction\. This shows that raw open\-source data contains a weak training signal that can be concentrated via dedicated curation, and that scaling*usable*tokens matters more than simply increasing token budgets\. The breakdown of our best open\-source configuration is detailed in Table[4](https://arxiv.org/html/2608.18655#A1.T4)\.
2. 2\.Synthetic data provides better quality and a higher volume\.Because it is generated at a larger scale, using a relatively high\-quality teacher model, its unfiltered sentence pairs already benefit from higher quality scores than raw open\-source pairs \(Figure[3](https://arxiv.org/html/2608.18655#S4.F3)\)\. This allows us to apply much stricterz¯\\mathrm\{\\bar\{z\}\}score thresholds while still retaining sufficient training data\. In competitive configurations, open\-source mixes typically require relatively permissive thresholds aroundz¯∈\[−0\.5,0\.5\]\\mathrm\{\\bar\{z\}\\in\[\-0\.5,0\.5\]\}, whereas synthetic mixes remain viable under substantially stricter thresholds, fromz¯≥0\.68\\mathrm\{\\bar\{z\}\\geq 0\.68\}toz¯≥1\.2\\mathrm\{\\bar\{z\}\\geq 1\.2\}\. As a result, synthetic mixes dominate open\-source mixes across nearly all training\-token budgets\.
3. 3\.Combining open\-source and synthetic mixes may help at smaller scales\.Combined mixes can outperform synthetic\-only mixes at smaller token budgets, but this advantage generally dematerialises as the budget is increased\. For example, the best large combined mix uses more tokens and a lower threshold than the final synthetic mix \(46\.49B tokens,z¯≥0\.41\\mathrm\{\\bar\{z\}\\geq 0\.41\}vs\. 32\.37B tokens,z¯≥0\.68\\mathrm\{\\bar\{z\}\\geq 0\.68\}\), yet performs slightly worse\. This suggests that once enough high\-quality synthetic data is available, adding lower\-quality open\-source data can dilute the training signal\.
4. 4\.TopN capping is associated with better token efficiency\.Incorporating topN with threshold\-only configurations as a means to limit the influence of well\-resourced language pairs tends to underperform in absolute performance although it tends to deliver a higher relative performance \(per token\)\. However, the effect is weak, therefore, restricted computational budgets should generally consider the topN capping option\.
5. 5\.Bidirectional expansion as a robust strategy\.Across open\-source, synthetic, and combined mixes, configurations with bidirectional expansion always improve compared to the equivalent configuration without expansion\. This suggests that expanding selected examples to both translation directions improves coverage and keeps the data more balanced\. Because filtering is still aligned with the final training direction, bidirectional expansion preserves QE reliability without the degradation caused by reversed filtering\.

Based on these observations, we choose the best\-performing synthetic mix \(threshold \+ bidirectional expansion\) as our final training dataset with 32\.37B tokens and reaching0\.632 SSA\-COMET\.

### 5\.3Which Dataset\(s\) Should We Use?

In this section, we evaluate several datasets in isolation to investigate their contribution to MT\.

![Refer to caption](https://arxiv.org/html/2608.18655v1/figures/single_source_bouquet_africa_iid_ssa_comet.png)

Figure 5:Single\-source SFT ablation on BOUQuET \(Qwen3\.5\-2B\)\.Each model is fine\-tuned using a single data component and evaluated with SSA\-COMET\. The percentages indicate relative gains over the baseline\.#### AfriNLLB

was the largest open\-source, curated dataset131313Refer to Appendix[A\.5](https://arxiv.org/html/2608.18655#A1.SS5)for dataset preparation\.to date hence we positioned it as our most competitive baseline in Figure[5](https://arxiv.org/html/2608.18655#S5.F5)\. While it is more beneficial than the human translations, its contribution only has a moderate effect on African MT\.

#### Human Mix

is our highest\-quality data, as it comes from projects with a high emphasis on human\-quality translations\. Figure[3](https://arxiv.org/html/2608.18655#S4.F3)shows that the Human Mix receives higher QE scores than AfriNLLB, which helps explain the fact that it achieves comparable performance despite being much smaller \(60\.2M versus 535M tokens\)\. This suggests that human\-quality data is highly valuable per token, but its limited scale makes it insufficient for effective SLM post\-training on its own\.

#### Instruct\-Mix

The primary reason for including this dataset in the final TranslatePsy\-AfriSLM mix is to preserve the conversational skills of our SLMs, see Figure[13](https://arxiv.org/html/2608.18655#A4.F13)for a short demo\. However, we observe that it unexpectedly benefits machine translation as well, even more so than a dedicated AfriNLLB corpus \(Figure[5](https://arxiv.org/html/2608.18655#S5.F5)\)\. This is almost certainly due to the 2\.3M examples \(around 50% of total\) of diverse tasks in African languages\. This effect is similar to the AfriqueLLM[51](https://arxiv.org/html/2608.18655#bib.bib5)continued pre\-training findings, which also observed an improvement in MT via instruction\-following SFT\.

#### TranslatePsy\-AfriSLM

Our data141414The best synthetic mix, highlighted in Figure[4](https://arxiv.org/html/2608.18655#S5.F4)\(top right\)\.emerges as a powerful contributor to African MT\. It combines scale and quality more effectively than prior sources: it retains enough data under strict filtering \(Figure[4](https://arxiv.org/html/2608.18655#S5.F4)\) and achieves higher QE scores than both AfriNLLB and the Human Mix \(Figure[3](https://arxiv.org/html/2608.18655#S4.F3)\)\. This makes it especially effective for African MT adaptation\. Detailed statistics are given in Table[5](https://arxiv.org/html/2608.18655#A1.T5)\.

### 5\.4Can SLMs Beat Frontier LLMs?

ModelFlores\-200BOUQuETSmolGeneral\-purpose LLMsApertus\-8B0\.38140\.40240\.3246Qwen3\.5\-27B0\.51320\.53360\.4279Qwen3\.5\-122B\-A10B0\.55050\.57160\.4574Dedicated translation modelsAfriNLLB\-600M0\.57290\.60380\.4748NLLB\-1\.3B0\.58780\.61300\.4850NLLB\-3\.3B0\.59440\.61780\.4909Hunyuan\-MT\-7B0\.44500\.44510\.3784TranslateGemma\-4B0\.39130\.40690\.3384TranslateGemma\-27B0\.54550\.56770\.4608African language\-specialized modelsAfriqueLlama\-8B0\.52100\.56200\.4423AfriqueGemma\-12B0\.55310\.58050\.4608AfriqueQwen\-14B0\.54430\.57180\.4527OursTranslatePsy\-AfriSLM\-0\.8B0\.59440\.62230\.4973TranslatePsy\-AfriSLM\-2B0\.60700\.63220\.5074TranslatePsy\-AfriSLM\-4B0\.61430\.63910\.5136

Table 3:TranslatePsy\-AfriSLM compared to LLMs and specialised translatorson Flores\-200, BOUQuET, Smol \(SSA\-COMET\)\. Full results in Table[17](https://arxiv.org/html/2608.18655#A4.T17)and[18](https://arxiv.org/html/2608.18655#A4.T18)\.In this section, we combine the TranslatePsy\-AfriSLM training dataset with the Instruct\-Mix and Asia\-Europe Mix to improve real\-world usability and contrast the performance of our SLMs with results from related methodologies\.

#### Africa\-IID

Table[3](https://arxiv.org/html/2608.18655#S5.T3)summarizes SSA\-COMET performance across three benchmarks; full results across all four metrics are reported in Tables[17](https://arxiv.org/html/2608.18655#A4.T17)and[18](https://arxiv.org/html/2608.18655#A4.T18)\(Appendix[D](https://arxiv.org/html/2608.18655#A4)\)\.\(1\) General\-purpose LLMsunderperform, even at frontier scale\. Qwen3\.5\-122B\-A10B is surpassed by our 0\.8B model on Flores\-200, BOUQuET, and Smol, suggesting that model scale alone is insufficient for African MT\.\(2\) Dedicated translation modelsalso lag behind our SLMs\. TranslatePsy\-AfriSLM\-0\.8B outperforms AfriNLLB\-600M, NLLB\-3\.3B and TranslateGemma\-27B across all benchmarks\. Notably, our 0\.8B model matches NLLB\-3\.3B on Flores\-200 and exceeds it on BOUQuET and Smol, despite using roughly a quarter of its parameters\. This demonstrates that our carefully curated post\-training data can produce powerful and efficient translation SLMs\.\(3\) African\-centric modelsnarrow the gap but still do not match our 0\.8B SLM\. AfriqueGemma\-12B and AfriqueQwen\-14B remain below TranslatePsy\-AfriSLM\-0\.8B on the reported benchmarks, while requiring substantially more compute at both training and inference time\.

#### Statistical significance\.

Paired bootstrap tests confirm that the main rankings in Table[3](https://arxiv.org/html/2608.18655#S5.T3)are statistically reliable \(Appendix[D\.2](https://arxiv.org/html/2608.18655#A4.SS2)\)\. The results show that TranslatePsy\-AfriSLM\-0\.8B significantly outperforms much larger LLM baselines on most settings, while TranslatePsy\-AfriSLM\-2B consistently surpasses dedicated NLLB models\. We also observe statistically significant gains from TranslatePsy\-AfriSLM\-0\.8B to 2B & 4B\.

![Refer to caption](https://arxiv.org/html/2608.18655v1/figures/africa_ood_per_language_ssa.png)Figure 6:Africa\-OOD performance \(SSA\-COMET\)\. Qwen3\.5\-2B versus TranslatePsy\-AfriSLM\-2B\. Scores are averaged over translation directions and all datasets\.
#### Africa\-OOD

We use Africa\-OOD to test whether training on our 19 target languages transfers to new languages\. Under SSA\-COMET, TranslatePsy\-AfriSLM\-2B improves over Qwen3\.5\-2B on all eight held\-out languages, with particularly large gains for the lower\-resource languages such as Sepedi, Bambara, and Akan \(Figure[6](https://arxiv.org/html/2608.18655#S5.F6)\)\. Further per\-language analyses show similar trends across COMET\-22, MetricX, and ChrF\+\+ \(Figure[11](https://arxiv.org/html/2608.18655#A3.F11); Appendix[D\.1](https://arxiv.org/html/2608.18655#A4.SS1)\)\. These results suggest meaningful cross\-lingual transfer beyond the training distribution, although less uniform than the IID gains\.

#### Catastrophic Forgetting

We include the Asia\-Europe Mix \(§[3\.4](https://arxiv.org/html/2608.18655#S3.SS4)\) to mitigate catastrophic forgetting on non\-African languages\. Figure[7](https://arxiv.org/html/2608.18655#S5.F7)shows that SFT without this data causes substantial degradation on Asian and European languages, most notably under MetricX \(−86\.0%\-86\.0\\%\)\. Adding our Asia\-Europe Mix reduces this to−10\.3%\-10\.3\\%and also mitigates degradation measured by COMET\-22, SSA\-COMET, and ChrF\+\+\. These results indicate that non\-African parallel data helps preserve broader multilingual translation ability without sacrificing Africa\-IID performance; in fact, we observe a small improvement on Africa\-IID pairs \(Appendix[C](https://arxiv.org/html/2608.18655#A3)\)\.

![Refer to caption](https://arxiv.org/html/2608.18655v1/figures/europe_asia_mix_forgetting_vs_qwen__bouquet__all_langs.png)Figure 7:Catastrophic forgetting on 38 Asian and European languages\.The bars show the percentage change from baseline after SFT, with/without the mix\.

## 6Conclusions

We have introducedTranslatePsy\-AfriSLM, a comprehensive investigation of training data curation aimed at improving machine translation for 19 Sub\-Saharan languages\. Our results lead to two main conclusions: 1\) Combining multiple QE metrics with robustz¯\\bar\{z\}score normalization, and scoring examples in the same direction as training, yields more consistent filtering than relying on any single metric\. This allows us to reduce the number of training tokens by up to 96% while maintaining comparable performance, 2\) Across the 1B–50B token range, filtered synthetic data dominates the quality\-efficiency Pareto frontier, while open\-source data appears to saturate early due to quality limitations\. Given the current shortage of high\-quality open\-source parallel data for African MT, filtered synthetic generation seems to be the most practical path forward\. Guided by these findings, we post\-trained TranslatePsy\-AfriSLM on our highest\-quality data, outperforming much larger models, including TranslateGemma\-27B and Qwen3\.5\-122B\-A10B, with as few as 0\.8B parameters\. We open\-source our data to support future work on making African MT a default capability in multilingual models\.

## Limitations

#### The Need for Human Evaluation\.

One notable finding that warrants future investigation is the challenge of determining absolute translation quality for African languages\. Using thereference\-free quality estimatorsfrom prior work, it appears that our TranslatePsy\-AfriSLM data is comparable to human translated pairs in quality but it comes in quantities that are orders of magnitude larger\. Still, the performance, as indicated by thereference\-based evaluation metrics\(produced by same model families as QE\) suggests that we have not yet reached the translation quality of European and Asian languages\. Therefore, conducting purely quantitative data scaling efforts using existing tools may run intounknown absolute performance limitations\. The reason\(s\) behind this can only be resolved by analysing each step in the pipeline \[data collection and training of QE models, correlation with human judgement, data collection and training of evaluation models, their correlation with ground truth\],leveraging expert human annotatorsto identify any errors that have the capacity to compound over multiple steps during a large\-scale data curation\.

## Ethics Statement

#### Data licensing

All primary sources are open\-source and publicly available under permissive attribution\-based licenses — ODC\-By v1\.0 \(Fine Translations, MADLAD\-400\) and CC BY 4\.0 \(MaLA Corpus\), alongside research\-permissive aggregates from OPUS and WMT22\. The compiled TranslatePsy\-AfriSLM datasets will be released under a similarly permissive license for safe, lawful, reproducible low\-resource NLP research\.

#### Dialectal representativeness

Exact dialect tracking is unavailable for this web\-curated corpus; models may over\-reflect standardized written forms and under\-represent regional dialects or oral traditions\. We advise auditing trained models for localized sensitivity before deployment\.

#### Synthetic content

Because large internet repositories involve data recycling, mixing, duplication, and some LLM\-based corrections and paraphrasing, even our "open\-source" subset should be treated as at least partially synthetic \(this is in addition to our purely synthetic translations\)\.

## References

- Adebaraet al\.\(2022\)I\. Adebara, A\. Elmadany, M\. Abdul\-Mageed, and A\. InciarteAfroLID: a neural language identification tool for African languages\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Y\. Goldberg, Z\. Kozareva, and Y\. Zhang \(Eds\.\),Abu Dhabi, United Arab Emirates,pp\. 1958–1981\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.128/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.128)Cited by:[item 3](https://arxiv.org/html/2608.18655#A1.I1.i3.p1.1)\.
- Adelaniet al\.\(2022a\)D\. I\. Adelani, J\. O\. Alabi, A\. Fan, J\. Kreutzer, X\. Shen, M\. Reid, D\. Ruiter, D\. Klakow, P\. Nabende, E\. Chang, T\. Gwadabe, F\. Sackey, B\. F\. P\. Dossou, C\. Emezue, C\. Leong, M\. Beukman, S\. H\. Muhammad, G\. D\. Jarso, O\. Yousuf, A\. N\. Niyongabo Rubungo, G\. Hacheme, E\. P\. Wairagala, M\. U\. Nasir, B\. A\. Ajibade, T\. O\. Ajayi, Y\. W\. Gitau, J\. Abbott, M\. Ahmed, M\. Ochieng, A\. Aremu, P\. Ogayo, J\. Mukiibi, F\. Ouoba Kabore, G\. K\. Kalipe, D\. Mbaye, A\. A\. Tapo, V\. M\. Memdjokam Koagne, E\. Munkoh\-Buabeng, V\. Wagner, I\. Abdulmumin, A\. Awokoya, H\. Buzaaba, B\. Sibanda, A\. Bukula, and S\. ManthaluA few thousand translations go a long way\! leveraging pre\-trained models for African news translation\.InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,M\. Carpuat, M\. de Marneffe, and I\. V\. Meza Ruiz \(Eds\.\),Seattle, United States,pp\. 3053–3070\.External Links:[Link](https://aclanthology.org/2022.naacl-main.223/),[Document](https://dx.doi.org/10.18653/v1/2022.naacl-main.223)Cited by:[§2\.1](https://arxiv.org/html/2608.18655#S2.SS1.SSS0.Px3.p1.1)\.
- Adelaniet al\.\(2022b\)D\. I\. Adelani, M\. M\. I\. Alam, A\. Anastasopoulos, A\. Bhagia, M\. R\. Costa\-jussà, J\. Dodge, F\. Faisal, C\. Federmann, N\. Fedorova, F\. Guzmán, S\. Koshelev, J\. Maillard, V\. Marivate, J\. Mbuya, A\. Mourachko, S\. Saleem, H\. Schwenk, and G\. WenzekFindings of the WMT’22 shared task on large\-scale machine translation evaluation for African languages\.InProceedings of the Seventh Conference on Machine Translation \(WMT\),P\. Koehn, L\. Barrault, O\. Bojar, F\. Bougares, R\. Chatterjee, M\. R\. Costa\-jussà, C\. Federmann, M\. Fishel, A\. Fraser, M\. Freitag, Y\. Graham, R\. Grundkiewicz, P\. Guzman, B\. Haddow, M\. Huck, A\. Jimeno Yepes, T\. Kocmi, A\. Martins, M\. Morishita, C\. Monz, M\. Nagata, T\. Nakazawa, M\. Negri, A\. Névéol, M\. Neves, M\. Popel, M\. Turchi, and M\. Zampieri \(Eds\.\),Abu Dhabi, United Arab Emirates \(Hybrid\),pp\. 773–800\.External Links:[Link](https://aclanthology.org/2022.wmt-1.72/),[Document](https://dx.doi.org/10.18653/v1/2022.wmt-1.72)Cited by:[§2\.1](https://arxiv.org/html/2608.18655#S2.SS1.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.18655#S3.SS1.SSS0.Px1.p1.1)\.
- Adelaniet al\.\(2023\)D\. I\. Adelani, M\. Masiak, I\. A\. Azime, J\. O\. Alabi, A\. L\. Tonja, C\. Mwase, O\. Ogundepo, B\. F\. P\. Dossou, A\. Oladipo, D\. Nixdorf, C\. C\. Emezue, S\. S\. al\-azzawi, B\. K\. Sibanda, D\. David, L\. Ndolela, J\. Mukiibi, T\. O\. Ajayi, T\. M\. Ngoli, B\. Odhiambo, A\. T\. Owodunni, N\. C\. Obiefuna, S\. H\. Muhammad, S\. S\. Abdullahi, M\. G\. Yigezu, T\. Gwadabe, I\. Abdulmumin, M\. T\. Bame, O\. O\. Awoyomi, I\. Shode, T\. A\. Adelani, H\. A\. Kailani, A\. Omotayo, A\. Adeeko, A\. Abeeb, A\. Aremu, O\. Samuel, C\. Siro, W\. Kimotho, O\. R\. Ogbu, C\. E\. Mbonu, C\. I\. Chukwuneke, S\. Fanijo, J\. Ojo, O\. F\. Awosan, T\. K\. Guge, S\. T\. Sari, P\. Nyatsine, F\. Sidume, O\. Yousuf, M\. Oduwole, U\. Kimanuka, K\. P\. Tshinu, T\. Diko, S\. Nxakama, A\. T\. Johar, S\. Gebre, M\. Mohamed, S\. A\. Mohamed, F\. M\. Hassan, M\. A\. Mehamed, E\. Ngabire, and P\. StenetorpMasakhaNEWS: news topic classification for african languages\.ArXiv\.Cited by:[§A\.3](https://arxiv.org/html/2608.18655#A1.SS3.SSS0.Px1.p1.1)\.
- Adelaniet al\.\(2025\)D\. I\. Adelani, J\. Ojo, I\. A\. Azime, J\. Y\. Zhuang, J\. Alabi, X\. He, M\. Ochieng, S\. Hooker, A\. Bukula, E\. A\. Lee,et al\.Irokobench: a new benchmark for african languages in the age of large language models\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 2732–2757\.Cited by:[§A\.3](https://arxiv.org/html/2608.18655#A1.SS3.SSS0.Px1.p1.1)\.
- Adelaniet al\.\(2021\)D\. Adelani, D\. Ruiter, J\. Alabi, D\. Adebonojo, A\. Ayeni, M\. Adeyemi, A\. E\. Awokoya, and C\. España\-BonetThe effect of domain and diacritics in Yoruba–English neural machine translation\.InProceedings of the 18th Biennial Machine Translation Summit \(Volume 1: Research Track\),Virtual,pp\. 61–75\.External Links:[Link](https://aclanthology.org/2021.mtsummit-research.6)Cited by:[§2\.1](https://arxiv.org/html/2608.18655#S2.SS1.SSS0.Px3.p1.1)\.
- Alabiet al\.\(2025\)J\. Alabi, I\. A\. Azime, M\. Zhang, C\. España\-Bonet, R\. Bawden, D\. Zhu, D\. I\. Adelani, C\. O\. Odoje, I\. Akinade, I\. Maab,et al\.AFRIDOC\-mt: document\-level mt corpus for african languages\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 27758–27794\.Cited by:[§A\.2](https://arxiv.org/html/2608.18655#A1.SS2.p1.1),[§2\.1](https://arxiv.org/html/2608.18655#S2.SS1.SSS0.Px3.p1.1)\.
- Bakouchet al\.\(2025\)E\. Bakouch, L\. Ben Allal, A\. Lozhkov, N\. Tazi, L\. Tunstall, C\. M\. Patiño, E\. Beeching, A\. Roucher, A\. J\. Reedi, Q\. Gallouédec, K\. Rasul, N\. Habib, C\. Fourrier, H\. Kydlicek, G\. Penedo, H\. Larcher, M\. Morlon, V\. Srivastav, J\. Lochner, X\. Nguyen, C\. Raffel, L\. von Werra, and T\. WolfSmolLM3: smol, multilingual, long\-context reasoner\.Note:[https://huggingface\.co/blog/smollm3](https://huggingface.co/blog/smollm3)Cited by:[§A\.3](https://arxiv.org/html/2608.18655#A1.SS3.SSS0.Px3.p1.1)\.
- Caswellet al\.\(2025\)I\. Caswell, E\. Nielsen, J\. Luo, C\. Cherry, G\. Kovacs, H\. Shemtov, P\. Talukdar, D\. Tewari, B\. M\. Diane, K\. M\. Doumbouya, D\. Diane, S\. F\. Cissé, E\. Ferrante, A\. Guasoni, M\. K\. Keita, S\. DebBarma, A\. Kuzhuget, D\. Anugraha, M\. R\. S\. Habibi, S\. Ahmadi, M\. Lau, and J\. EngSMOL: Professionally translated parallel data for 115 under\-represented languages\.External Links:2502\.12301,[Link](https://arxiv.org/abs/2502.12301)Cited by:[§A\.2](https://arxiv.org/html/2608.18655#A1.SS2.p1.1),[§2\.1](https://arxiv.org/html/2608.18655#S2.SS1.SSS0.Px3.p1.1),[§4\.2](https://arxiv.org/html/2608.18655#S4.SS2.SSS0.Px2.p1.1)\.
- Deutschet al\.\(2025\)D\. Deutsch, E\. Briakou, I\. Caswell, M\. Finkelstein, R\. Galor, J\. Juraska, G\. Kovacs, A\. Lui, R\. Rei, J\. Riesa, S\. Rijhwani, P\. Riley, E\. Salesky, F\. Trabelsi, S\. Winkler, B\. Zhang, and M\. FreitagWMT24\+\+: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects\.External Links:2502\.12404,[Link](https://arxiv.org/abs/2502.12404)Cited by:[§2\.1](https://arxiv.org/html/2608.18655#S2.SS1.SSS0.Px3.p1.1)\.
- Devineet al\.\(2026\)P\. Devine, M\. Sanni, F\. Adilazuarda, J\. G\. Loizaga, and B\. HaddowKakugo: distillation of low\-resource languages into small language models\.arXiv preprint arXiv:2601\.14051\.Cited by:[§A\.3](https://arxiv.org/html/2608.18655#A1.SS3.SSS0.Px1.p1.1)\.
- Dialloet al\.\(2025\)K\. Diallo, J\. Smith, C\. T\. Okolo, D\. Nyamwaya, J\. Kgomo, and R\. NgamitaCase studies of ai policy development in africa\.Data & Policy7,pp\. e15\.Cited by:[§1](https://arxiv.org/html/2608.18655#S1.p1.1)\.
- Dinget al\.\(2023\)N\. Ding, Y\. Chen, B\. Xu, Y\. Qin, Z\. Zheng, S\. Hu, Z\. Liu, M\. Sun, and B\. ZhouEnhancing chat language models by scaling high\-quality instructional conversations\.arXiv preprint arXiv:2305\.14233\.Cited by:[§A\.3](https://arxiv.org/html/2608.18655#A1.SS3.SSS0.Px1.p1.1)\.
- Emezue and Dossou \(2021\)C\. C\. Emezue and B\. F\. P\. DossouMMTAfrica: multilingual machine translation for African languages\.InProceedings of the Sixth Conference on Machine Translation,L\. Barrault, O\. Bojar, F\. Bougares, R\. Chatterjee, M\. R\. Costa\-jussa, C\. Federmann, M\. Fishel, A\. Fraser, M\. Freitag, Y\. Graham, R\. Grundkiewicz, P\. Guzman, B\. Haddow, M\. Huck, A\. J\. Yepes, P\. Koehn, T\. Kocmi, A\. Martins, M\. Morishita, and C\. Monz \(Eds\.\),Online,pp\. 398–411\.External Links:[Link](https://aclanthology.org/2021.wmt-1.48/)Cited by:[§2\.1](https://arxiv.org/html/2608.18655#S2.SS1.SSS0.Px3.p1.1)\.
- Finkelsteinet al\.\(2026\)M\. Finkelstein, I\. Caswell, T\. Domhan, J\. Peter, J\. Juraska, P\. Riley, D\. Deutsch, C\. Dilanni, C\. Cherry, E\. Briakou,et al\.TranslateGemma technical report\.arXiv preprint arXiv:2601\.09012\.Cited by:[§1](https://arxiv.org/html/2608.18655#S1.p1.1),[§4\.2](https://arxiv.org/html/2608.18655#S4.SS2.SSS0.Px4.p1.1)\.
- Hasanet al\.\(2021\)T\. Hasan, A\. Bhattacharjee, Md\. S\. Islam, K\. Mubasshir, Y\. Li, Y\. Kang, M\. S\. Rahman, and R\. ShahriyarXL\-sum: large\-scale multilingual abstractive summarization for 44 languages\.InFindings of the Association for Computational Linguistics: ACL\-IJCNLP 2021,Online,pp\. 4693–4703\.External Links:[Link](https://aclanthology.org/2021.findings-acl.413)Cited by:[§A\.3](https://arxiv.org/html/2608.18655#A1.SS3.SSS0.Px1.p1.1)\.
- Hernández\-Canoet al\.\(2025\)A\. Hernández\-Cano, A\. Hägele, A\. H\. Huang, A\. Romanou, A\. Solergibert, B\. Pasztor, B\. Messmer, D\. Garbaya, E\. F\. Ďurech, I\. Hakimi, J\. G\. Giraldo, M\. Ismayilzada, N\. Foroutan, S\. Moalla, T\. Chen, V\. Sabolčec, Y\. Xu, M\. Aerni, B\. AlKhamissi, I\. A\. Marinas, M\. H\. Amani, M\. Ansaripour, I\. Badanin, H\. Benoit, E\. Boros, N\. Browning, F\. Bösch, M\. Böther, N\. Canova, C\. Challier, C\. Charmillot, J\. Coles, J\. Deriu, A\. Devos, L\. Drescher, D\. Dzenhaliou, M\. Ehrmann, D\. Fan, S\. Fan, S\. Gao, M\. Gila, M\. Grandury, D\. Hashemi, A\. Hoyle, J\. Jiang, M\. Klein, A\. Kucharavy, A\. Kucherenko, F\. Lübeck, R\. Machacek, T\. Manitaras, A\. Marfurt, K\. Matoba, S\. Matrenok, H\. Mendoncça, F\. R\. Mohamed, S\. Montariol, L\. Mouchel, S\. Najem\-Meyer, J\. Ni, G\. Oliva, M\. Pagliardini, E\. Palme, A\. Panferov, L\. Paoletti, M\. Passerini, I\. Pavlov, A\. Poiroux, K\. Ponkshe, N\. Ranchin, J\. Rando, M\. Sauser, J\. Saydaliev, M\. A\. Sayfiddinov, M\. Schneider, S\. Schuppli, M\. Scialanga, A\. Semenov, K\. Shridhar, R\. Singhal, A\. Sotnikova, A\. Sternfeld, A\. K\. Tarun, P\. Teiletche, J\. Vamvas, X\. Yao, H\. Z\. A\. Ilic, A\. Klimovic, A\. Krause, C\. Gulcehre, D\. Rosenthal, E\. Ash, F\. Tramèr, J\. VandeVondele, L\. Veraldi, M\. Rajman, T\. Schulthess, T\. Hoefler, A\. Bosselut, M\. Jaggi, and I\. SchlagApertus: Democratizing Open and Compliant LLMs for Global Language Environments\.Note:[https://arxiv\.org/abs/2509\.14233](https://arxiv.org/abs/2509.14233)Cited by:[§1](https://arxiv.org/html/2608.18655#S1.p1.1),[§4\.2](https://arxiv.org/html/2608.18655#S4.SS2.SSS0.Px4.p1.1)\.
- Isangula \(2025\)K\. G\. IsangulaNavigating barriers: challenges and strategies for adopting artificial intelligence in qualitative research in low\-income african contexts\.Tanzania Journal of Health Research25\(3\),pp\. 2048–2059\.Cited by:[§1](https://arxiv.org/html/2608.18655#S1.p1.1)\.
- Jiet al\.\(2024\)S\. Ji, Z\. Li, I\. Paul, J\. Paavola, P\. Lin, P\. Chen, D\. O’Brien, H\. Luo, H\. Schütze, J\. Tiedemann, and B\. HaddowEMMA\-500: enhancing massively multilingual adaptation of large language models\.arXiv preprint 2409\.17892\.External Links:[Link](https://arxiv.org/abs/2409.17892)Cited by:[§2\.1](https://arxiv.org/html/2608.18655#S2.SS1.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.18655#S3.SS1.SSS0.Px1.p1.1)\.
- Joulinet al\.\(2016\)A\. Joulin, E\. Grave, P\. Bojanowski, and T\. MikolovBag of tricks for efficient text classification\.arXiv preprint arXiv:1607\.01759\.Cited by:[item 3](https://arxiv.org/html/2608.18655#A1.I1.i3.p1.1)\.
- Juraskaet al\.\(2024\)J\. Juraska, D\. Deutsch, M\. Finkelstein, and M\. FreitagMetricX\-24: the Google submission to the WMT 2024 metrics shared task\.InProceedings of the Ninth Conference on Machine Translation,B\. Haddow, T\. Kocmi, P\. Koehn, and C\. Monz \(Eds\.\),Miami, Florida, USA,pp\. 492–504\.External Links:[Link](https://aclanthology.org/2024.wmt-1.35)Cited by:[§4\.2](https://arxiv.org/html/2608.18655#S4.SS2.SSS0.Px3.p1.1)\.
- Juraskaet al\.\(2023\)J\. Juraska, M\. Finkelstein, D\. Deutsch, A\. Siddhant, M\. Mirzazadeh, and M\. FreitagMetricX\-23: The Google Submission to the WMT 2023 Metrics Shared Task\.InProceedings of the Eighth Conference on Machine Translation,P\. Koehn, B\. Haddow, T\. Kocmi, and C\. Monz \(Eds\.\),Singapore,pp\. 756–767\.External Links:[Link](https://aclanthology.org/2023.wmt-1.63),[Document](https://dx.doi.org/10.18653/v1/2023.wmt-1.63)Cited by:[§2\.2](https://arxiv.org/html/2608.18655#S2.SS2.p1.1)\.
- Kuduguntaet al\.\(2023\)S\. Kudugunta, I\. Caswell, B\. Zhang, X\. Garcia, C\. A\. Choquette\-Choo, K\. Lee, D\. Xin, A\. Kusupati, R\. Stella, A\. Bapna, and O\. FiratMADLAD\-400: a multilingual and document\-level large audited dataset\.External Links:2309\.04662Cited by:[§3\.1](https://arxiv.org/html/2608.18655#S3.SS1.SSS0.Px2.p1.1)\.
- Liet al\.\(2025\)S\. Li, J\. Wang, F\. D\. M\. A\. Ali, C\. Cherry, D\. Deutsch, E\. Briakou, R\. Sousa\-Silva, H\. Lopes Cardoso, P\. Stenetorp, and D\. I\. AdelaniSSA\-COMET: do LLMs outperform learned metrics in evaluating MT for under\-resourced African languages?\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 12979–12998\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.656/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.656),ISBN 979\-8\-89176\-332\-6Cited by:[§2\.2](https://arxiv.org/html/2608.18655#S2.SS2.p1.1),[§4\.2](https://arxiv.org/html/2608.18655#S4.SS2.SSS0.Px3.p1.1),[footnote 4](https://arxiv.org/html/2608.18655#footnote4)\.
- Maluleke \(2025\)A\. F\. MalulekeAI adoption in african higher education: a systematic review of benefits and ethical implications\.Interdisciplinary Journal of Education Research7\(2\),pp\. a05–a05\.Cited by:[§1](https://arxiv.org/html/2608.18655#S1.p1.1)\.
- Moslemet al\.\(2026\)Y\. Moslem, A\. K\. Wassie, and A\. G\. AbebeAfriNLLB: efficient translation models for african languages\.External Links:[Link](https://api.semanticscholar.org/CorpusID:285463140)Cited by:[§2\.1](https://arxiv.org/html/2608.18655#S2.SS1.SSS0.Px2.p1.1),[§2\.2](https://arxiv.org/html/2608.18655#S2.SS2.p1.1),[§4\.2](https://arxiv.org/html/2608.18655#S4.SS2.SSS0.Px4.p1.1)\.
- Moukatib and Seddik \(2026\)M\. Moukatib and A\. B\. SeddikThe the role of ai in translator training: assessing ai’s influence on translation education and professional training\.International Journal of Linguistics and Translation Studies7\(1\),pp\. 136–151\.Cited by:[§1](https://arxiv.org/html/2608.18655#S1.p1.1)\.
- Muhammadet al\.\(2023\)S\. H\. Muhammad, I\. Abdulmumin, S\. M\. Yimam, D\. I\. Adelani, I\. S\. Ahmad, N\. Ousidhoum, A\. Ayele, S\. M\. Mohammad, and M\. BeloucifSemEval\-2023 task 12: sentiment analysis for african languages \(afrisenti\-semeval\)\.arXiv preprint arXiv:2304\.06845\.Cited by:[§A\.3](https://arxiv.org/html/2608.18655#A1.SS3.SSS0.Px1.p1.1)\.
- NVIDIA \(2024\)NVIDIANeMo curator: gpu\-accelerated data curation for training ai models\.External Links:[Link](https://github.com/NVIDIA-NeMo/Curator)Cited by:[§A\.1](https://arxiv.org/html/2608.18655#A1.SS1.p1.1)\.
- Nwagbalaet al\.\(2025\)S\. C\. Nwagbala, F\. N\. Ezeanokwasa, R\. Nwachukwu, N\. J\. Uzodike, and O\. P\. NwosuAI adoption and sustainability of smes in africa: opportunities and challenges\.International Journal of Science and Research Archive14\(1\),pp\. 467–475\.Cited by:[§1](https://arxiv.org/html/2608.18655#S1.p1.1)\.
- Ogundepoet al\.\(2023\)O\. Ogundepo, T\. R\. Gwadabe, C\. E\. Rivera, J\. H\. Clark, S\. Ruder, D\. I\. Adelani, B\. F\. P\. Dossou, A\. A\. DIOP, C\. Sikasote, G\. Hacheme, H\. Buzaaba, I\. Ezeani, R\. Mabuya, S\. Osei, C\. Emezue, A\. N\. Kahira, S\. H\. Muhammad, A\. Oladipo, A\. T\. Owodunni, A\. L\. Tonja, I\. Shode, A\. Asai, T\. O\. Ajayi, C\. Siro, S\. Arthur, M\. Adeyemi, O\. Ahia, A\. Anuoluwapo, O\. Awosan, C\. Chukwuneke, B\. Opoku, A\. Ayodele, V\. Otiende, C\. Mwase, B\. Sinkala, A\. N\. Rubungo, D\. A\. Ajisafe, E\. F\. Onwuegbuzia, H\. Mbow, E\. Niyomutabazi, E\. Mukonde, F\. I\. Lawan, I\. S\. Ahmad, J\. O\. Alabi, M\. Namukombo, M\. Chinedu, M\. Phiri, N\. Putini, N\. Mngoma, P\. A\. Amuok, R\. N\. Iro, and S\. AdhiamboAfriQA: cross\-lingual open\-retrieval question answering for african languages\.External Links:2305\.06897Cited by:[§A\.3](https://arxiv.org/html/2608.18655#A1.SS3.SSS0.Px1.p1.1)\.
- Ojoet al\.\(2025\)J\. Ojo, O\. Ogundepo, A\. Oladipo, K\. Ogueji, J\. Lin, P\. Stenetorp, and D\. I\. AdelaniAfroBench: how good are large language models on african languages?\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 19048–19095\.Cited by:[§A\.3](https://arxiv.org/html/2608.18655#A1.SS3.SSS0.Px1.p1.1)\.
- Olmoet al\.\(2025\)T\. Olmo, A\. Ettinger, A\. Bertsch, B\. Kuehl, D\. Graham, D\. Heineman, D\. Groeneveld, F\. Brahman, F\. Timbers, H\. Ivison, J\. Morrison, J\. Poznanski, K\. Lo, L\. Soldaini, M\. Jordan, M\. Chen, M\. Noukhovitch, N\. Lambert, P\. Walsh, P\. Dasigi, R\. Berry, S\. Malik, S\. Shah, S\. Geng, S\. Arora, S\. Gupta, T\. Anderson, T\. Xiao, T\. Murray, T\. Romero, V\. Graf, A\. Asai, A\. Bhagia, A\. Wettig, A\. Liu, A\. Rangapur, C\. Anastasiades, C\. Huang, D\. Schwenk, H\. Trivedi, I\. Magnusson, J\. Lochner, J\. Liu, L\. J\. V\. Miranda, M\. Sap, M\. Morgan, M\. Schmitz, M\. Guerquin, M\. Wilson, R\. Huff, R\. L\. Bras, R\. Xin, R\. Shao, S\. Skjonsberg, S\. Z\. Shen, S\. S\. Li, T\. Wilde, V\. Pyatkin, W\. Merrill, Y\. Chang, Y\. Gu, Z\. Zeng, A\. Sabharwal, L\. Zettlemoyer, P\. W\. Koh, A\. Farhadi, N\. A\. Smith, and H\. HajishirziOlmo 3\.External Links:2512\.13961,[Link](https://arxiv.org/abs/2512.13961)Cited by:[§A\.3](https://arxiv.org/html/2608.18655#A1.SS3.SSS0.Px3.p1.1)\.
- Owusu \(2025\)M\. OwusuAfri code datasets collection\.Hugging Face\.Note:A collection of datasets featuring coding conversations translated into 55 African languagesExternal Links:[Link](https://huggingface.co/collections/michsethowusu/afri-code-datasets)Cited by:[§A\.3](https://arxiv.org/html/2608.18655#A1.SS3.SSS0.Px1.p1.1)\.
- Penedoet al\.\(2026\)G\. Penedo, H\. Kydlíček, A\. H\. Kargaran, and L\. von WerraFineTranslations\.Hugging Face\.Note:[https://huggingface\.co/datasets/HuggingFaceFW/finetranslations](https://huggingface.co/datasets/HuggingFaceFW/finetranslations)Cited by:[§2\.1](https://arxiv.org/html/2608.18655#S2.SS1.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.18655#S3.SS1.SSS0.Px1.p1.1)\.
- Popović \(2017\)M\. PopovićChrF\+\+: words helping character n\-grams\.InProceedings of the Second Conference on Machine Translation,O\. Bojar, C\. Buck, R\. Chatterjee, C\. Federmann, Y\. Graham, B\. Haddow, M\. Huck, A\. J\. Yepes, P\. Koehn, and J\. Kreutzer \(Eds\.\),Copenhagen, Denmark,pp\. 612–618\.External Links:[Link](https://aclanthology.org/W17-4770/),[Document](https://dx.doi.org/10.18653/v1/W17-4770)Cited by:[§4\.2](https://arxiv.org/html/2608.18655#S4.SS2.SSS0.Px3.p1.1)\.
- Reiet al\.\(2022a\)R\. Rei, J\. G\. C\. de Souza, D\. Alves, C\. Zerva, A\. C\. Farinha, T\. Glushkova, A\. Lavie, L\. Coheur, and A\. F\. T\. MartinsCOMET\-22: unbabel\-IST 2022 submission for the metrics shared task\.InProceedings of the Seventh Conference on Machine Translation \(WMT\),P\. Koehn, L\. Barrault, O\. Bojar, F\. Bougares, R\. Chatterjee, M\. R\. Costa\-jussà, C\. Federmann, M\. Fishel, A\. Fraser, M\. Freitag, Y\. Graham, R\. Grundkiewicz, P\. Guzman, B\. Haddow, M\. Huck, A\. Jimeno Yepes, T\. Kocmi, A\. Martins, M\. Morishita, C\. Monz, M\. Nagata, T\. Nakazawa, M\. Negri, A\. Névéol, M\. Neves, M\. Popel, M\. Turchi, and M\. Zampieri \(Eds\.\),Abu Dhabi, United Arab Emirates \(Hybrid\),pp\. 578–585\.External Links:[Link](https://aclanthology.org/2022.wmt-1.52/),[Document](https://dx.doi.org/10.18653/v1/2022.wmt-1.52)Cited by:[§4\.2](https://arxiv.org/html/2608.18655#S4.SS2.SSS0.Px3.p1.1)\.
- Reiet al\.\(2020\)R\. Rei, C\. Stewart, A\. C\. Farinha, and A\. LavieCOMET: a neural framework for MT evaluation\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 2685–2702\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.213/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.213)Cited by:[§2\.2](https://arxiv.org/html/2608.18655#S2.SS2.p1.1)\.
- Reiet al\.\(2022b\)R\. Rei, M\. Treviso, N\. M\. Guerreiro, C\. Zerva, A\. C\. Farinha, C\. Maroti, J\. G\. C\. de Souza, T\. Glushkova, D\. Alves, L\. Coheur, A\. Lavie, and A\. F\. T\. MartinsCometKiwi: IST\-unbabel 2022 submission for the quality estimation shared task\.InProceedings of the Seventh Conference on Machine Translation \(WMT\),P\. Koehn, L\. Barrault, O\. Bojar, F\. Bougares, R\. Chatterjee, M\. R\. Costa\-jussà, C\. Federmann, M\. Fishel, A\. Fraser, M\. Freitag, Y\. Graham, R\. Grundkiewicz, P\. Guzman, B\. Haddow, M\. Huck, A\. Jimeno Yepes, T\. Kocmi, A\. Martins, M\. Morishita, C\. Monz, M\. Nagata, T\. Nakazawa, M\. Negri, A\. Névéol, M\. Neves, M\. Popel, M\. Turchi, and M\. Zampieri \(Eds\.\),Abu Dhabi, United Arab Emirates \(Hybrid\),pp\. 634–645\.External Links:[Link](https://aclanthology.org/2022.wmt-1.60/),[Document](https://dx.doi.org/10.18653/v1/2022.wmt-1.60)Cited by:[item 1](https://arxiv.org/html/2608.18655#A1.I5.i1.p1.1),[§2\.2](https://arxiv.org/html/2608.18655#S2.SS2.p1.1)\.
- Sadvilkar and Neumann \(2020\)N\. Sadvilkar and M\. NeumannPySBD: pragmatic sentence boundary disambiguation\.InProceedings of Second Workshop for NLP Open Source Software \(NLP\-OSS\),Online,pp\. 110–114\.External Links:[Link](https://www.aclweb.org/anthology/2020.nlposs-1.15)Cited by:[item 4](https://arxiv.org/html/2608.18655#A1.I1.i4.p1.1)\.
- Singhet al\.\(2024\)S\. Singh, F\. Vargus, D\. Dsouza, B\. F\. Karlsson, A\. Mahendiran, W\. Ko, H\. Shandilya, J\. Patel, D\. Mataciunas, L\. OMahony, M\. Zhang, R\. Hettiarachchi, J\. Wilson, M\. Machado, L\. S\. Moura, D\. Krzemiński, H\. Fadaei, I\. Ergün, I\. Okoh, A\. Alaagib, O\. Mudannayake, Z\. Alyafeai, V\. M\. Chien, S\. Ruder, S\. Guthikonda, E\. A\. Alghamdi, S\. Gehrmann, N\. Muennighoff, M\. Bartolo, J\. Kreutzer, A\. Üstün, M\. Fadaee, and S\. HookerAya dataset: an open\-access collection for multilingual instruction tuning\.External Links:2402\.06619Cited by:[§A\.3](https://arxiv.org/html/2608.18655#A1.SS3.SSS0.Px1.p1.1)\.
- Ssemugabi \(2025\)S\. SsemugabiThe role of ai in modern language translation and its societal applications: a systematic literature review\.InSouthern African Conference for Artificial Intelligence Research,pp\. 390–404\.Cited by:[§1](https://arxiv.org/html/2608.18655#S1.p1.1)\.
- Taoriet al\.\(2023\)R\. Taori, I\. Gulrajani, T\. Zhang, Y\. Dubois, X\. Li, C\. Guestrin, P\. Liang, and T\. B\. HashimotoStanford alpaca: an instruction\-following llama model\.GitHub\.Note:[https://github\.com/tatsu\-lab/stanford\_alpaca](https://github.com/tatsu-lab/stanford_alpaca)Cited by:[§A\.3](https://arxiv.org/html/2608.18655#A1.SS3.SSS0.Px1.p1.1)\.
- Teamet al\.\(2022\)N\. Team, M\. R\. Costa\-jussà, J\. Cross, O\. Çelebi, M\. Elbayad, K\. Heafield, K\. Heffernan, E\. Kalbassi, J\. Lam, D\. Licht, J\. Maillard, A\. Sun, S\. Wang, G\. Wenzek, A\. Youngblood, B\. Akula, L\. Barrault, G\. M\. Gonzalez, P\. Hansanti, J\. Hoffman, S\. Jarrett, K\. R\. Sadagopan, D\. Rowe, S\. Spruit, C\. Tran, P\. Andrews, N\. F\. Ayan, S\. Bhosale, S\. Edunov, A\. Fan, C\. Gao, V\. Goswami, F\. Guzmán, P\. Koehn, A\. Mourachko, C\. Ropers, S\. Saleem, H\. Schwenk, and J\. WangNo language left behind: scaling human\-centered machine translation\.arXiv preprint arXiv: 2207\.04672\.Cited by:[§2\.1](https://arxiv.org/html/2608.18655#S2.SS1.SSS0.Px1.p1.1),[§4\.2](https://arxiv.org/html/2608.18655#S4.SS2.SSS0.Px2.p1.1),[§4\.2](https://arxiv.org/html/2608.18655#S4.SS2.SSS0.Px4.p1.1)\.
- Team \(2026\)Q\. TeamQwen3\.5: accelerating productivity with native multimodal agents\.External Links:[Link](https://qwen.ai/blog?id=qwen3.5)Cited by:[§1](https://arxiv.org/html/2608.18655#S1.p1.1),[§4\.1](https://arxiv.org/html/2608.18655#S4.SS1.p1.1),[§4\.2](https://arxiv.org/html/2608.18655#S4.SS2.SSS0.Px4.p1.1)\.
- Teamet al\.\(2026\)T\. O\. M\. Team, B\. Alastruey, N\. Bafna, A\. Caciolai, K\. Heffernan, A\. Kozhevnikov, C\. Ropers, E\. Sánchez, C\. Saint\-James, I\. Tsiamas, C\. Cheng, J\. Chuang, P\. Duquenne, M\. Duppenthaler, N\. Ekberg, C\. Gao, P\. L\. H\. Cabot, J\. M\. Janeiro, J\. Maillard, G\. M\. Gonzalez, H\. Schwenk, E\. Toledo, A\. Turkatenko, A\. Ventayol\-Boada, R\. Moritz, A\. Mourachko, S\. Parimi, M\. Williamson, S\. Yates, D\. Dale, and M\. R\. Costa\-jussàOmnilingual MT: machine translation for 1,600 languages\.External Links:,[Link](https://arxiv.org/abs/2603.16309)Cited by:[§4\.2](https://arxiv.org/html/2608.18655#S4.SS2.SSS0.Px2.p1.1)\.
- Tiedemann \(2012\)J\. TiedemannParallel data, tools and interfaces in OPUS\.InProceedings of the Eighth International Conference on Language Resources and Evaluation \(LREC’12\),N\. Calzolari, K\. Choukri, T\. Declerck, M\. U\. Doğan, B\. Maegaard, J\. Mariani, A\. Moreno, J\. Odijk, and S\. Piperidis \(Eds\.\),Istanbul, Turkey,pp\. 2214–2218\.External Links:[Link](http://www.lrec-conf.org/proceedings/lrec2012/pdf/463_Paper.pdf)Cited by:[§A\.4](https://arxiv.org/html/2608.18655#A1.SS4.p1.1),[§2\.1](https://arxiv.org/html/2608.18655#S2.SS1.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.18655#S3.SS1.SSS0.Px1.p1.1)\.
- Uemuraet al\.\(2026\)K\. Uemura, M\. Zhang, and D\. I\. AdelaniAfriMTEB and AfriE5: benchmarking and adapting text embedding models for African languages\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 3697–3717\.External Links:[Link](https://aclanthology.org/2026.eacl-long.171/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.171)Cited by:[§2\.2](https://arxiv.org/html/2608.18655#S2.SS2.p1.1)\.
- Wanget al\.\(2024\)J\. Wang, D\. I\. Adelani, S\. Agrawal, M\. Masiak, R\. Rei, E\. Briakou, M\. Carpuat, X\. He, S\. Bourhim, A\. Bukula,et al\.AfriMTE and africomet: enhancing comet to embrace under\-resourced african languages\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 5997–6023\.Cited by:[§2\.2](https://arxiv.org/html/2608.18655#S2.SS2.p1.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§1](https://arxiv.org/html/2608.18655#S1.p1.1),[§4\.2](https://arxiv.org/html/2608.18655#S4.SS2.SSS0.Px4.p1.1)\.
- Yuet al\.\(2026\)H\. Yu, T\. Xu, M\. A\. Hedderich, W\. Hamidouche, S\. W\. Zamir, and D\. I\. AdelaniAfriqueLLM: how data mixing and model architecture impact continued pre\-training for african languages\.arXiv preprint arXiv:2601\.06395\.Cited by:[§2\.1](https://arxiv.org/html/2608.18655#S2.SS1.SSS0.Px2.p1.1),[§2\.2](https://arxiv.org/html/2608.18655#S2.SS2.p1.1),[§4\.2](https://arxiv.org/html/2608.18655#S4.SS2.SSS0.Px4.p1.1),[§5\.3](https://arxiv.org/html/2608.18655#S5.SS3.SSS0.Px3.p1.1)\.
- Zhanget al\.\(2020\)B\. Zhang, P\. Williams, I\. Titov, and R\. SennrichImproving massively multilingual neural machine translation and zero\-shot translation\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 1628–1639\.External Links:[Link](https://aclanthology.org/2020.acl-main.148),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.148)Cited by:[§A\.4](https://arxiv.org/html/2608.18655#A1.SS4.p1.1)\.
- Zhenget al\.\(2025\)M\. Zheng, Z\. Li, B\. Qu, M\. Song, Y\. Du, M\. Sun, and D\. WangHunyuan\-mt technical report\.arXiv preprint arXiv:2509\.05209\.Cited by:[§1](https://arxiv.org/html/2608.18655#S1.p1.1),[§4\.2](https://arxiv.org/html/2608.18655#S4.SS2.SSS0.Px4.p1.1)\.

## Appendix AAppendix

### A\.1Synthetic and Open\-Source Mixes

The corpora composing these mixes are described in §[3\.1](https://arxiv.org/html/2608.18655#S3.SS1)\. Here, we give more details of the processing steps and final data statistics\. The preprocessing steps \(up to decontamination\) were implemented with NeMo Curator[29](https://arxiv.org/html/2608.18655#bib.bib50)\. In Tables[4](https://arxiv.org/html/2608.18655#A1.T4)and[5](https://arxiv.org/html/2608.18655#A1.T5)\), "Clean" means before deduplication, "Preproc" after decontamination, "Filtered" after QE filtering, and "Bidir\." after bidirectional expansion\.

#### Processing steps:

1. 1\.Document Splitting:151515Document and sentence splitting was performed only on monolingual source texts to be translated for synthetic parallel data generation, because the teacher is not accurate/designed for paragraph\-long input texts\.has divided documents into paragraphs \(splitting at newlines\)\.
2. 2\.Cleaning:removed Unicode artifacts, extra newlines, markup, and URLs\.
3. 3\.Language Identification:viaAfroLID[1](https://arxiv.org/html/2608.18655#bib.bib36)for African languages andFastText[20](https://arxiv.org/html/2608.18655#bib.bib51)for English\.
4. 4\.Sentence Splitting:00footnotemark:0was performed using pySBD[40](https://arxiv.org/html/2608.18655#bib.bib52)\.
5. 5\.Filtering:removed non\-alphanumeric content and boilerplate text\.
6. 6\.Deduplication:with exact & fuzzy matching\.
7. 7\.Eval\-set Decontamination:removed sentences similar to our test sets \(§[A\.6](https://arxiv.org/html/2608.18655#A1.SS6)\)\.
8. 8\.Synthetic Data GenerationThe monolingual data \(§[3\.1](https://arxiv.org/html/2608.18655#S3.SS1)\) was translated with the NLLB\-3\.3B model,161616[https://huggingface\.co/facebook/nllb\-200\-3\.3B](https://huggingface.co/facebook/nllb-200-3.3B)decoded with the c\-translate2 decoder171717[https://github\.com/OpenNMT/CTranslate2](https://github.com/OpenNMT/CTranslate2)with "int8\_float16" quantisation and beam size 3\. We discarded the 54B MoE model because the MoE architecture is not supported by fast decoders\.
9. 9\.Quality Estimation \(QE\):For open\-source data, the input to the QE stage is the parallel pre\-processed text\. For the synthetic data, it is the monolingual pre\-processed source text together with its translation\. For each pair, MetricX, AfriCOMET, SSA\-COMET scores and the z\-score were calculated in the source\-target and target\-source directions\. Aligned filtering was applied \(§[3\.2](https://arxiv.org/html/2608.18655#S3.SS2)\), and pairs with a z\-score below the threshold were discarded\.

QE FilteredLang\.RawCleanPreprocen→\\toxxxx→\\toenxx→\\toyyBidir\.afr66,08948,93243,0601,5001,0002,2487,425amh41,10936,93530,5421,0207762,8416,518hau33,51524,38022,5371,230421,4534,162ibo17,55015,37814,3851,500173093,458kin21,85516,42515,4841,50002802,989lin9,9926,1695,7271,500102,217lug14,2578,1757,5281,3024801,922mlg11,1068,3367,8291,39410192,477nya14,65510,7159,804586145371,952orm11,9495,7035,517892871,175sna21,47712,67611,7441,391115103,488som26,15722,43220,0681,388392923,425sot4,5173,6763,491231055swa37,66430,78428,2231,500416975,947tsn14,5959,3258,5201,60925633,085wol4,3423,5063,2808681141,218xho22,97114,43513,5341,500132003,595yor13,9099,0528,09141167691,999zul24,47613,70912,9501,1321533,196others15,00414,82814,2921,5001,00005,000Total427,187315,571286,60623,7473,0859,62365,304

Table 4:Open\-Source mix statisticsby language \(line counts in thousands\)\.MonolingualParallel, QE FilteredLang\.RawCleanPreproceng→\\toxxxx→\\toengBidir\.eng145,726145,688136,482———afr24,06223,38411,1391,0091,8756,974amh3,7043,5763,3794,89756910,786hau5,0024,7294,4206,6971,30617,261ibo2,4502,3512,2225,96463912,659kin5,2114,6994,2192,5368478,102lin1531451355,056367,482lug3783583325,5864810,089mlg3,5212,6902,4641,0104704,168nya2,2921,8461,7683,5754398,060orm7570684,817107,580sna4263443352,490925,829som6,9816,6165,9672,7448517,381sot1,9801,8511,7544,5765038,495swa18,08017,55915,8343,5982,77617,670tsn2221203,50676,033wol45433913,256723,615xho2,2192,0401,9154,4144657,962yor2,3452,1092,00713,4641,12827,140zul2,1181,7371,6629,11551918,367Total226,790221,856196,16298,31212,589215,653

Table 5:Synthetic\-Mix stats\(TranslatePsy\-AfriSLM training data\) by language \(line counts in thousands\)\.Quality estimatorC22SSAMXChrF\+\+eng\-xx \(English→\\rightarrowAfrican language\(s\)\)z\-score0\.7370\.6174\.5245\.7SSA\-COMET0\.7360\.6204\.7345\.9AfriCOMET0\.7350\.6094\.8745\.3MetricX0\.7300\.6054\.2945\.2xx\-eng \(African language\(s\)→\\rightarrowEnglish\)z\-score0\.7540\.5815\.3052\.0SSA\-COMET0\.7460\.5815\.7951\.9AfriCOMET0\.7560\.5795\.4052\.3MetricX0\.7490\.5735\.2751\.7Table 6:Translation quality with individual versus unified QE metricsover Flores\-200 test set\.Quality estimatorC22SSAMXChrF\+\+eng\-xx \(English→\\rightarrowAfrican language\(s\)\)z\-score0\.6550\.5038\.8828\.0SSA\-COMET0\.6520\.5059\.2228\.0AfriCOMET0\.6530\.4959\.2627\.8MetricX0\.6500\.4918\.5827\.9xx\-eng \(African language\(s\)→\\rightarrowEnglish\)z\-score0\.5700\.50010\.2529\.1SSA\-COMET0\.5650\.49810\.6529\.2AfriCOMET0\.5700\.49710\.4029\.3MetricX0\.5680\.49110\.1429\.2Table 7:Translation quality with individual versus unified QE metricsover Smol test set\.

### A\.2Human Mix

We compiled 352,582 human\-translated, high\-quality parallel sentences focused on African languages from AfriDOC\-MT[7](https://arxiv.org/html/2608.18655#bib.bib12)and SMOL181818Excluding our evaluation set \(smolsent\)\.[9](https://arxiv.org/html/2608.18655#bib.bib11)\. Table[8](https://arxiv.org/html/2608.18655#A1.T8)summarizes the high\-level statistics\. The SmolDoc subset of SMOL191919[https://huggingface\.co/datasets/google/smol](https://huggingface.co/datasets/google/smol)was flattened to sentence\-level pairs\.

#### Processing Steps:

1. 1\.Bidirectional expansion:Both directions \(en→\\rightarrowxx and xx→\\rightarrowen\) were used for training\.
2. 2\.Global exact deduplication:Identical pairs were removed\.
3. 3\.Eval\-set decontamination:Removed sentences similar to our evaluation data\.
4. 4\.Global approximate deduplication:MinHash applied to remove similar examples\.

Filter TypeKeptRemovedRetained %Raw sentence pairs180,540–100\.00%\(1\) After bidirectional expansion361,080–200\.00%\(2\) After global exact deduplication357,5843,49699\.03%\(3\) After eval\-set decontamination357,40218299\.95%\(4\) After global approximate deduplication352,5824,82098\.65%

Table 8:Filtering summary for the human translations\.

### A\.3Instruct Mix

#### Afri\-Instruct Mix

This dataset is a heterogeneous mixture of open\-ended instruction data aggregated from 11 publicly available HuggingFace datasets[41](https://arxiv.org/html/2608.18655#bib.bib22);[43](https://arxiv.org/html/2608.18655#bib.bib23);[11](https://arxiv.org/html/2608.18655#bib.bib24);[34](https://arxiv.org/html/2608.18655#bib.bib48);[28](https://arxiv.org/html/2608.18655#bib.bib25);[4](https://arxiv.org/html/2608.18655#bib.bib26);[31](https://arxiv.org/html/2608.18655#bib.bib29);[5](https://arxiv.org/html/2608.18655#bib.bib30);[32](https://arxiv.org/html/2608.18655#bib.bib31);[16](https://arxiv.org/html/2608.18655#bib.bib27);[13](https://arxiv.org/html/2608.18655#bib.bib32)\. Table[9](https://arxiv.org/html/2608.18655#A1.T9)summarizes each data source\. The final mixture contains 2,306,800 examples designed to maintain/extend essential instruction\-following abilities to African MT, e\.g\. open\-ended instruction and chat data, which contributes 1,553,944 examples \(67\.36%\), followed by code\-assistance data with 521,389 examples \(22\.60%\)\. Smaller portions come from classification and other structured prediction tasks with 123,448 examples \(5\.35%\), summarization and headline generation with 74,581 examples \(3\.23%\), and question answering plus cross\-lingual QA with 33,438 examples \(1\.45%\)\.

#### Processing Steps:

1. 1\.Format standardization:Each dataset was converted to a multi\-turn chat format with system, user, and assistant roles\. Appropriate system prompts were added \(e\.g\., “You are a sentiment classifier” for AfriSenti, “You are a summarization assistant” for XLSum\), etc\.
2. 2\.Quality filtering:Samples with empty user or assistant messages were removed\.
3. 3\.Task expansion:AfriQA was expanded into three task types: machine translation pairs, monolingual QA, and cross\-lingual QA, tripling its effective sample count\.
4. 4\.Sampling:For Afri\-Code datasets, a maximum of 30,000 samples per language were randomly sampled to reduce repetition\.

DatasetSourceCountLanguagesAfrican\-UltraChatmasakhane/african\-ultrachat54,99411African\-Alpacamasakhane/african\-translated\-alpaca832,02916Kakugoptrdvn/kakugo\-\{lang\}464,56912Aya DatasetCohereLabs/aya\_dataset202,35265Afri\-Codemichsethowusu/Code\-170k\-\{lang\}521,38918AfriSentishmuhammad/AfriSenti\-twitter\-sentiment83,6888MasakhaNEWSmasakhane/masakhanews21,29612XLSumcsebuetnlp/xlsum53,2857AfriADRmasakhane/AfriADR26,1103AfriXNLImasakhane/afrixnli13,65013AfriQAmasakhane/afriqa33,4387Total\-2,306,800\-

Table 9:Data sources for the Afri\-Instruct data mix\.
#### General\-Instruct Mix

Additional instruction\-following training examples \(mostly English\-centric\) come from two public HuggingFace datasets:smoltalk2, the post\-training data of SmolLM3[8](https://arxiv.org/html/2608.18655#bib.bib33)andDolci\-Instruct, the OLMO\-3 supervised fine\-tuning data[33](https://arxiv.org/html/2608.18655#bib.bib34)\. Table[10](https://arxiv.org/html/2608.18655#A1.T10)summarizes each data source\. Exact deduplication removed 85,982 examples\. Due to the exclusion of reasoning data in TranslatePsy\-AfriSLM training, it does not retain the ’thinking’ capability of its base model\.

#### Processing Steps:

1. 1\.Format standardization:Datasets were converted to a unified chat format\. A system prompt \(“You are a helpful assistant\.”\) was added to conversations from Dolci\-Instruct\.
2. 2\.Exact deduplication:Identical examples were removed\.

DatasetSourceSplitsSmolTalk2HuggingFaceTB/smoltalk2\(SFT\)multilingual\_8languages\_lang\_5\_no\_thinksmollm3\_systemchats\_30k\_no\_thinksmollm3\_everyday\_conversations\_no\_thinksmollm3\_explore\_instruct\_rewriting\_no\_thinksmollm3\_smol\_rewrite\_no\_thinksmollm3\_smol\_summarize\_no\_thinkDolci\-Instructallenai/Dolci\-Instruct\-SFT\-No\-ToolstrainTotal\-2,308,569

Table 10:Data sources for the General Instruct mix\.

### A\.4Asia\-Europe Mix

In order to mitigate catastrophic forgetting, we included a broad selection of 38 medium\-high resource languages202020[https://huggingface\.co/datasets/Helsinki\-NLP/opus\-100](https://huggingface.co/datasets/Helsinki-NLP/opus-100)from OPUS\-100[52](https://arxiv.org/html/2608.18655#bib.bib13);[47](https://arxiv.org/html/2608.18655#bib.bib14)\. The final dataset comprises 24,114,303 parallel sentences\. We only use en→\\rightarrowxx pairs as the xx→\\rightarrowen direction is robust to large\-scale African post\-training\. Including this data has negligible impact on African translation quality, while it plays a crucial role in preserving Asian and European performance \(Appendix[C](https://arxiv.org/html/2608.18655#A3)\)\. Table[11](https://arxiv.org/html/2608.18655#A1.T11)provides an aggregate summary while Table[12](https://arxiv.org/html/2608.18655#A1.T12)in summarizes the per\-language statistics\.

#### Processing Steps:

1. 1\.
2. 2\.Approximate deduplication:MinHash was applied using 128 permutations and a Jaccard similarity of 0\.8 on character 4\-grams\.
3. 3\.Eval\-set decontamination:Sentences similar to our evaluation sets were removed\.
4. 4\.Global approximate deduplication:BPE\-unigram MinHash deduplication \(256 permutations, Jaccard threshold 0\.8\) was applied globally to remove similar examples\.

Filter TypeKeptRemovedRetained%Raw sentence pairs36,347,268–100\.00%After COMET mean≥0\.6\\geq 0\.630,112,8766,234,39282\.85%After pair\-level \(local\) MinHash deduplication25,439,4664,673,41084\.48%After eval\-set decontamination25,429,12910,33799\.96%After MinHash deduplication24,114,3031,314,82694\.83%

Table 11:Filtering summary for Europe/Asia mix\.PairRAWCOMETDEDUPPairRAWCOMETDEDUPen\-zh1,000k837k767ken\-sv1,000k828k725ken\-ja1,000k788k669ken\-hr1,000k829k738ken\-it1,000k843k752ken\-is1,000k650k485ken\-ru1,000k817k751ken\-lt1,000k879k650ken\-es1,000k879k791ken\-ms1,000k815k634ken\-tr1,000k880k778ken\-id1,000k860k720ken\-fr1,000k839k789ken\-si979k829k475ken\-pl1,000k787k695ken\-hi534k474k298ken\-ar1,000k809k751ken\-bn1,000k896k604ken\-uk1,000k803k578ken\-ur754k727k603ken\-pt1,000k848k748ken\-fa1,000k813k730ken\-sk1,000k860k716ken\-kk80k57k41ken\-de1,000k773k720ken\-ro1,000k831k725ken\-hu1,000k828k732ken\-bg1,000k814k718ken\-el1,000k822k735ken\-cs1,000k814k720ken\-ko1,000k726k617ken\-da1,000k820k712ken\-vi1,000k863k729ken\-lv1,000k879k652ken\-th1,000k841k721ken\-nl1,000k830k747ken\-fi1,000k810k732ken\-et1,000k814k691k

Table 12:Language statistics for Asia\-Europe data\.

### A\.5AfriNLLB

We apply only test data decontamination and minimal approximate deduplication, see Table[13](https://arxiv.org/html/2608.18655#A1.T13)\.

#### Processing Steps:

1. 1\.Approximate deduplication:MinHash was applied using 128 permutations and a Jaccard similarity of 0\.8 on character 4\-grams\.
2. 2\.Eval\-set decontamination:Sentences similar to our evaluation sets were removed\.

Filter TypeKeptRemovedRetained%Original AfriNLLB pairs3,218,822–100\.00%After eval\-set decontamination3,206,91811,90499\.63%After approximate deduplication3,129,17577,74397\.58%

Table 13:AfriNLLB preprocessing steps\.

### A\.6Data Decontamination

As an essential part of our preprocessing steps, we decontaminate our training data against all evaluation sentences, languages and dataset splits, totalingover 850K sentences\. We used a BPE\-unigram \(Qwen3\.5\) MinHash \(LSH\) deduplication with a Jaccard similarity threshold of 0\.9 to filter training data on an individual sentence level rather than a pair level for the most granular detection possible\.

## Appendix BAdditional Experimental Setup

MT Prompt TemplateSystem Prompt:
You are a professionalSOURCE\_LANGtoTARGET\_LANGtranslator\. Your goal is to accurately convey the meaning and nuances of the originalSOURCE\_LANGtext while adhering toTARGET\_LANGgrammar, vocabulary, and cultural sensitivities\. Produce only theTARGET\_LANGtranslation, without any additional explanations or commentary\.User:
Please translate the followingSOURCE\_LANGtext intoTARGET\_LANG:SOURCE\_TEXT\. Translation:Assistant:
TARGET\_TEXTFigure 8:Training and evaluation prompt template\.### B\.1Translation Prompt

Training and evaluation examples were structured as shown in Figure[8](https://arxiv.org/html/2608.18655#A2.F8)whereSOURCE\_LANGandTARGET\_LANGdenote the source and target languagenameswhileSOURCE\_TEXTandTARGET\_TEXTrepresent the input/outputtextsin each language, respectively\. Finally, a model\-specific chat template was applied\.

### B\.2Training Hyperparameters

#### Optimization\.

Full fine\-tuning uses peak learning rate of1\.25×10−51\.25\\times 10^\{\-5\}; fused AdamW, a linear schedule with 1% warmup, gradient clipping at 1\.0, and gradient checkpointing, with a global batch size of 256\.

#### Data formatting\.

All training samples are first converted into the translation prompt format shown in Figure[8](https://arxiv.org/html/2608.18655#A2.F8), and then wrapped with the model\-specific chat template\. Sequences exceeding 2,048 tokens are filtered, and remaining samples are packed via best\-fit decreasing\.

#### Infrastructure\.

All experiments use PyTorch, HuggingFace Transformers, TRL, and DeepSpeed ZeRO\-2 in bfloat16 on 32 NVIDIA H100 GPUs\.

![Refer to caption](https://arxiv.org/html/2608.18655v1/figures/europe_asia_mix_forgetting_vs_qwen__bouquet.png)Figure 9:Effect of the Asia\-Europe Mix on catastrophic forgetting and Africa\-IID performance\.Bars show the percentage change from the Qwen3\.5\-2B baseline after SFT\. The Asia\-Europe Mix substantially reduces degradation on Asian and European language pairs while preserving nearly identical gains on the target Africa\-IID pairs\.

## Appendix CCatastrophic Forgetting Analysis

We further analyze whether the Asia\-Europe Mix mitigates catastrophic forgetting on non\-African languages without compromising African MT performance\. The mix contains bidirectional parallel data between English and non\-African languages, covering both Asian222222Asian: sin, hin, ben, urd, fas, kaz, tur, zho, jpn, kor, vie, tha, ind, msa, fil\.and European232323European: fra, spa, ita, por, ron, ell, lit, lav, rus, ukr, pol, slk, hrv, ces, srp, bul, deu, swe, isl, dan, nld, hun, fin, est\.language groups\.

Figure[9](https://arxiv.org/html/2608.18655#A2.F9)shows that fine\-tuning without the Asia\-Europe Mix substantially degrades performance on Asian and European language pairs, indicating catastrophic forgetting outside the African training distribution\. Adding the mix consistently reduces this degradation across all metrics and both language groups\. For example, the MetricX\-24 drop is reduced from−95\.1%\-95\.1\\%to−16\.8%\-16\.8\\%for Asian languages and from−81\.0%\-81\.0\\%to−6\.7%\-6\.7\\%for European languages\. This retention does not come at the cost of Africa\-IID performance: models with and without the Asia\-Europe Mix achieve nearly identical gains on African language pairs across all metrics\. We therefore retain the Asia\-Europe Mix in the final TranslatePsy\-AfriSLM recipe, not to improve African MT directly, but to preserve broader multilingual translation while maintaining the gains on the Africa\-IID group\.

![Refer to caption](https://arxiv.org/html/2608.18655v1/figures/africa_iid_bouquet_per_language.png)Figure 10:AFRICA\-IID BOUQuET results\.Per\-language scores for 19 target African languages on BOUQuET, comparing Qwen3\.5\-2B and TranslatePsy\-AfriSLM\-2B across four evaluation metrics\. Scores are averaged over both translation directions where available\.Δ\\Delta= TranslatePsy\-AfriSLM\-2B−\-baseline model scores\.![Refer to caption](https://arxiv.org/html/2608.18655v1/figures/africa_ood_per_language.png)Figure 11:AFRICA\-OOD results across all evaluation metrics\.Per\-language scores for held\-out African languages, comparing Qwen3\.5\-2B and TranslatePsy\-AfriSLM\-2B\. Scores are averaged over both translation directions and over the available evaluation sets among Flores\-200, BOUQuET, and Smol; language coverage differs by evaluation set as shown in Table[14](https://arxiv.org/html/2608.18655#A4.T14)\.Δ\\Delta= TranslatePsy\-AfriSLM\-2B−\-baseline model scores\.![Refer to caption](https://arxiv.org/html/2608.18655v1/figures/training_tokens_vs_metrics_africa_all__all_datasets.png)Figure 12:Scaling runs as QE quality and training budgets increase, evaluated on Flores\-200, BOUQuET, and Smol \(columns\)\. Each row reports a different metric, each point corresponds to a model trained with a particular data source and filtering configuration\. Colours distinguish filtered, synthetic, and filtered\+synthetic data, while marker shapes indicate the filtering setup\. MetricX is lower\-is\-better; all other metrics are higher\-is\-better\. Filtering thresholds are represented by an increasing fill, i\.e\. the fuller the shape, the higher the threshold\.
## Appendix DAdditional Results

AFRICA\-OOD lang\.Flores\-200BOUQuETSmolNigerian Pidgin \(pcm\)✓✓Sudanese Arabic \(apd\)✓Akan \(aka\)✓✓✓Tamazight \(ber\)✓✓✓Kituba \(ktu\)✓Bambara \(bam\)✓✓✓Sepedi \(nso\)✓✓✓Mooré \(mos\)✓✓✓\# Langs\.568\# Directions101216

Table 14:AFRICA\-OOD language coverage by evaluation set\.The number of directions counts both eng\-to\-xx and xx\-to\-eng directions\.### D\.1Per\-Language Performance Analysis

#### Africa\-IID

Figure[10](https://arxiv.org/html/2608.18655#A3.F10)shows per\-language results on the Africa\-IID BOUQuET benchmark, averaging scores over both eng\-xx and xx\-eng directions\. TranslatePsy\-AfriSLM\-2B consistently improves over the Qwen3\.5\-2B baseline across the 19 in\-distribution languages and all four evaluation metrics\. The largest gains appear for low\-baseline languages such as Oromo, Malagasy, Lingala, Tswana, and Zulu, while Afrikaans shows smaller gains due to its stronger baseline\.

#### Africa\-OOD

Figure[11](https://arxiv.org/html/2608.18655#A3.F11)reports results for Africa\-OOD languages held out from fine\-tuning\. Because Flores\-200, BOUQuET, and Smol cover different subsets of OOD languages, scores are averaged over both directions and available evaluation sets, as summarized in Table[14](https://arxiv.org/html/2608.18655#A4.T14)\. Although the OOD languages are held out from fine\-tuning, TranslatePsy\-AfriSLM\-2B still improves overall performance, suggesting meaningful transfer beyond the languages seen during post\-training\. The strongest gains appear for Sepedi, Bambara, Akan, and Mooré, while improvements are smaller and more heterogeneous than in the IID setting, with some metric\-specific regressions for Nigerian Pidgin, Sudanese Arabic, and Tamazight\.

Data SourceConfigurationThresholdTokensThresholdTokensThresholdTokensOpen\-Sourcethreshold\-0\.506\.57B0\.003\.79B0\.501\.76Bthreshold \+ top\-NN\-0\.285\.28B0\.162\.88B0\.611\.20Bthreshold \+ bidirectional\-0\.5014\.12B0\.007\.86B0\.503\.40Bthreshold \+ top\-NN\+ bidirectional\-0\.289\.95B0\.165\.15B0\.612\.03BSyntheticthreshold0\.6816\.75B0\.937\.28B1\.192\.80Bthreshold \+ top\-NN0\.7112\.43B0\.934\.49B1\.181\.44Bthreshold \+ bidirectional0\.6832\.37B0\.9315\.05B1\.196\.16Bthreshold \+ top\-NN\+ bidirectional0\.7123\.62B0\.939\.09B1\.183\.32BOpen\-Source \+ Syntheticthreshold0\.4123\.32B0\.6711\.08B0\.974\.56Bthreshold \+ top\-NN0\.5517\.71B0\.787\.37B1\.072\.64Bthreshold \+ bidirectional0\.4146\.49B0\.6722\.92B0\.979\.56Bthreshold \+ top\-NN\+ bidirectional0\.5533\.57B0\.7814\.24B1\.075\.35B

Table 15:Threshold values and training token counts for each data source, configuration, and filtering level\(for the datasets plotted in Figure[4](https://arxiv.org/html/2608.18655#S5.F4)\)\. Since the threshold may vary depending on the data mix and language direction, and topN changes the effective threshold, the value shown is an average weighted by the number of example pairs at each actual threshold\. The bolded configuration indicates the TranslatePsy\-AfriSLM data mix used for final training\.

### D\.2Statistical Significance \(Bootstrap\)

We assess statistical significance using paired bootstrap resampling over sentence pairs\. For each model pair, dataset, and metric, we align common translation directions, pair sentence\-level scores, and drawB=10,000B=10\{,\}000bootstrap samples with replacement\. We reportΔ\\Deltaas the mean score difference and computeppas the fraction of bootstrap samples that do not support the observed direction of improvement\. For MetricX, lower scores are better; for all other metrics, higher scores are better\.

Table[16](https://arxiv.org/html/2608.18655#A4.T16)shows that the main rankings are highly stable across benchmarks and metrics\. TranslatePsy\-AfriSLM\-0\.8B substantially outperforms the much larger general\-purpose Qwen3\.5\-122B on SSA\-COMET, MetricX, and ChrF\+\+, and remains competitive on COMET\-22\. It also consistently surpasses TranslateGemma\-27B and AfriGemma\-12B across all reported settings\. Against NLLB, TranslatePsy\-AfriSLM\-0\.8B is competitive but mixed, whereas TranslatePsy\-AfriSLM\-2B significantly outperforms both NLLB\-1\.3B and NLLB\-3\.3B across every benchmark–metric combination\. Within the TranslatePsy\-AfriSLM family, scaling from 0\.8B to 2B and from 2B to 4B yields monotonic and statistically significant gains\. Overall, the bootstrap results confirm that the observed improvements are stable under paired sentence\-level resampling rather than being driven by evaluation noise\.

FLORES\-200BOUQuETSMOLComparisonMetricΔ\\DeltappΔ\\DeltappΔ\\Deltappvs\. General\-purpose LLMTranslatePsy\-AfriSLM\-0\.8B vs\. Qwen3\.5\-122BCOMET\-22\-0\.00010\.452\+0\.0135<<0\.001\+0\.0017<<0\.001TranslatePsy\-AfriSLM\-0\.8B vs\. Qwen3\.5\-122BSSA\-COMET\+0\.0413<<0\.001\+0\.0379<<0\.001\+0\.0286<<0\.001TranslatePsy\-AfriSLM\-0\.8B vs\. Qwen3\.5\-122BMetricX\-0\.548<<0\.001\-0\.723<<0\.001\-0\.506<<0\.001TranslatePsy\-AfriSLM\-0\.8B vs\. Qwen3\.5\-122BChrF\+\+\+3\.30<<0\.001\+5\.33<<0\.001\+1\.29<<0\.001vs\. Dedicated TranslationTranslatePsy\-AfriSLM\-0\.8B vs\. NLLB\-1\.3BCOMET\-22\-0\.0039<<0\.001\+0\.00110\.017\+0\.0021<<0\.001TranslatePsy\-AfriSLM\-0\.8B vs\. NLLB\-1\.3BSSA\-COMET\+0\.0066<<0\.001\+0\.0086<<0\.001\+0\.0122<<0\.001TranslatePsy\-AfriSLM\-0\.8B vs\. NLLB\-1\.3BMetricX\+0\.0380\.002\-0\.154<<0\.001\-0\.242<<0\.001TranslatePsy\-AfriSLM\-0\.8B vs\. NLLB\-1\.3BChrF\+\+\-0\.40<<0\.001\+0\.48<<0\.001\-0\.010\.395TranslatePsy\-AfriSLM\-0\.8B vs\. NLLB\-3\.3BCOMET\-22\-0\.0093<<0\.001\+0\.0064<<0\.001\+0\.00110\.002TranslatePsy\-AfriSLM\-0\.8B vs\. NLLB\-3\.3BSSA\-COMET\+0\.00000\.483\+0\.0133<<0\.001\+0\.0124<<0\.001TranslatePsy\-AfriSLM\-0\.8B vs\. NLLB\-3\.3BMetricX\+0\.284<<0\.001\-0\.0070\.314\+0\.0040\.385TranslatePsy\-AfriSLM\-0\.8B vs\. NLLB\-3\.3BChrF\+\+\-1\.45<<0\.001\+0\.120\.101\-0\.19<<0\.001TranslatePsy\-AfriSLM\-2B vs\. NLLB\-1\.3BCOMET\-22\+0\.0081<<0\.001\+0\.0119<<0\.001\+0\.0091<<0\.001TranslatePsy\-AfriSLM\-2B vs\. NLLB\-1\.3BSSA\-COMET\+0\.0193<<0\.001\+0\.0189<<0\.001\+0\.0227<<0\.001TranslatePsy\-AfriSLM\-2B vs\. NLLB\-1\.3BMetricX\-0\.522<<0\.001\-0\.565<<0\.001\-0\.624<<0\.001TranslatePsy\-AfriSLM\-2B vs\. NLLB\-1\.3BChrF\+\+\+1\.28<<0\.001\+2\.42<<0\.001\+0\.75<<0\.001TranslatePsy\-AfriSLM\-2B vs\. NLLB\-3\.3BCOMET\-22\+0\.0027<<0\.001\+0\.0172<<0\.001\+0\.0081<<0\.001TranslatePsy\-AfriSLM\-2B vs\. NLLB\-3\.3BSSA\-COMET\+0\.0126<<0\.001\+0\.0237<<0\.001\+0\.0229<<0\.001TranslatePsy\-AfriSLM\-2B vs\. NLLB\-3\.3BMetricX\-0\.276<<0\.001\-0\.419<<0\.001\-0\.378<<0\.001TranslatePsy\-AfriSLM\-2B vs\. NLLB\-3\.3BChrF\+\+\+0\.24<<0\.001\+2\.06<<0\.001\+0\.57<<0\.001TranslatePsy\-AfriSLM\-0\.8B vs\. TGemma\-27BCOMET\-22\+0\.0097<<0\.001\+0\.0272<<0\.001\+0\.0018<<0\.001TranslatePsy\-AfriSLM\-0\.8B vs\. TGemma\-27BSSA\-COMET\+0\.0489<<0\.001\+0\.0547<<0\.001\+0\.0365<<0\.001TranslatePsy\-AfriSLM\-0\.8B vs\. TGemma\-27BMetricX\-0\.945<<0\.001\-1\.190<<0\.001\-0\.643<<0\.001TranslatePsy\-AfriSLM\-0\.8B vs\. TGemma\-27BChrF\+\+\+5\.08<<0\.001\+7\.79<<0\.001\+2\.33<<0\.001vs\. African\-SpecializedTranslatePsy\-AfriSLM\-0\.8B vs\. AfriGemma\-12BCOMET\-22\+0\.0301<<0\.001\+0\.0204<<0\.001\+0\.0212<<0\.001TranslatePsy\-AfriSLM\-0\.8B vs\. AfriGemma\-12BSSA\-COMET\+0\.0414<<0\.001\+0\.0306<<0\.001\+0\.0284<<0\.001TranslatePsy\-AfriSLM\-0\.8B vs\. AfriGemma\-12BMetricX\-0\.664<<0\.001\-0\.531<<0\.001\-0\.444<<0\.001TranslatePsy\-AfriSLM\-0\.8B vs\. AfriGemma\-12BChrF\+\+\+3\.57<<0\.001\+3\.19<<0\.001\+2\.03<<0\.001Scaling EffectTranslatePsy\-AfriSLM\-4B vs\. TranslatePsy\-AfriSLM\-2BCOMET\-22\+0\.0077<<0\.001\+0\.0080<<0\.001\+0\.0059<<0\.001TranslatePsy\-AfriSLM\-4B vs\. TranslatePsy\-AfriSLM\-2BSSA\-COMET\+0\.0073<<0\.001\+0\.0071<<0\.001\+0\.0063<<0\.001TranslatePsy\-AfriSLM\-4B vs\. TranslatePsy\-AfriSLM\-2BMetricX\-0\.320<<0\.001\-0\.287<<0\.001\-0\.260<<0\.001TranslatePsy\-AfriSLM\-4B vs\. TranslatePsy\-AfriSLM\-2BChrF\+\+\+1\.14<<0\.001\+1\.33<<0\.001\+0\.65<<0\.001TranslatePsy\-AfriSLM\-2B vs\. TranslatePsy\-AfriSLM\-0\.8BCOMET\-22\+0\.0120<<0\.001\+0\.0116<<0\.001\+0\.0085<<0\.001TranslatePsy\-AfriSLM\-2B vs\. TranslatePsy\-AfriSLM\-0\.8BSSA\-COMET\+0\.0126<<0\.001\+0\.0163<<0\.001\+0\.0152<<0\.001TranslatePsy\-AfriSLM\-2B vs\. TranslatePsy\-AfriSLM\-0\.8BMetricX\-0\.560<<0\.001\-0\.552<<0\.001\-0\.441<<0\.001TranslatePsy\-AfriSLM\-2B vs\. TranslatePsy\-AfriSLM\-0\.8BChrF\+\+\+1\.68<<0\.001\+1\.96<<0\.001\+0\.83<<0\.001TranslatePsy\-AfriSLM\-4B vs\. TranslatePsy\-AfriSLM\-0\.8BCOMET\-22\+0\.0197<<0\.001\+0\.0196<<0\.001\+0\.0143<<0\.001TranslatePsy\-AfriSLM\-4B vs\. TranslatePsy\-AfriSLM\-0\.8BSSA\-COMET\+0\.0199<<0\.001\+0\.0234<<0\.001\+0\.0214<<0\.001TranslatePsy\-AfriSLM\-4B vs\. TranslatePsy\-AfriSLM\-0\.8BMetricX\-0\.880<<0\.001\-0\.840<<0\.001\-0\.701<<0\.001TranslatePsy\-AfriSLM\-4B vs\. TranslatePsy\-AfriSLM\-0\.8BChrF\+\+\+2\.82<<0\.001\+3\.29<<0\.001\+1\.49<<0\.001

Table 16:Paired bootstrap significance tests\(B=10,000B=10\{,\}000\)\.Δ\\Delta= score difference \(ours−\-baseline\)\. Baseline comparisons use 0\.8B models\. BlueΔ\\Delta= ours better; red = ours worse \(MetricX: lower is better, negativeΔ\\Deltais blue\)\.ModelFlores\-200BOUQuETSmolSSA \(↑\)MX \(↓\)SSA \(↑\)MX \(↓\)SSA \(↑\)MX \(↓\)General\-purpose LLMs/SLMsApertus\-8B0\.381411\.4710\.402410\.3850\.324614\.233Qwen3\-4B0\.242615\.4270\.256514\.7720\.211717\.273Qwen3\-8B0\.265515\.3270\.276014\.7250\.227517\.336Qwen3\.5\-0\.8B0\.296715\.9740\.332014\.9020\.245417\.595Qwen3\.5\-2B0\.288514\.3440\.304913\.3170\.249815\.926Qwen3\.5\-4B0\.386311\.4160\.398410\.5670\.330113\.868Qwen3\.5\-9B0\.45239\.7580\.46648\.9270\.381113\.216Qwen3\.5\-27B0\.51327\.2500\.53366\.4680\.427911\.420Qwen3\.5\-122B\-A10B0\.55055\.9200\.57165\.3180\.457410\.515Dedicated translation modelsAfriNLLB\-600M0\.57295\.9160\.60384\.8310\.474810\.476NLLB\-600M0\.57185\.7520\.60194\.8030\.472310\.430NLLB\-1\.3B0\.58785\.1590\.61304\.4480\.485010\.036NLLB\-3\.3B0\.59444\.9130\.61784\.2640\.49099\.841Hunyuan\-MT\-7B0\.445010\.7790\.445110\.6590\.378413\.949TranslateGemma\-4B0\.391311\.0860\.406910\.2190\.338414\.000TranslateGemma\-12B0\.50377\.7180\.52286\.9260\.429911\.501TranslateGemma\-27B0\.54556\.1420\.56775\.5060\.460810\.468African language\-specialized modelsAfriqueLlama\-8B0\.52107\.0570\.56205\.6790\.442310\.974AfriqueGemma\-4B0\.52766\.6920\.55875\.7160\.441010\.938AfriqueGemma\-12B0\.55315\.8620\.58055\.0930\.460810\.397AfriqueQwen\-8B0\.52226\.8710\.55495\.7480\.439410\.959AfriqueQwen\-14B0\.54436\.2430\.57185\.4080\.452710\.748OursTranslatePsy\-AfriSLM\-0\.8B0\.59445\.1970\.62234\.3160\.49739\.823TranslatePsy\-AfriSLM\-2B0\.60704\.6370\.63223\.9400\.50749\.458TranslatePsy\-AfriSLM\-4B0\.61434\.3170\.63913\.7010\.51369\.232Table 17:SSA\-COMET and MetricXresults on Flores\-200, BOUQuET, and Smol\.ModelFlores\-200BOUQuETSmolC22 \(↑\)ChrF\+\+ \(↑\)C22 \(↑\)ChrF\+\+ \(↑\)C22 \(↑\)ChrF\+\+ \(↑\)General\-purpose LLMsApertus\-8B0\.588330\.450\.602527\.840\.491218\.02Qwen3\-4B0\.404617\.340\.430514\.250\.360710\.74Qwen3\-8B0\.456620\.640\.479217\.290\.400212\.50Qwen3\.5\-0\.8B0\.365614\.020\.392011\.170\.32558\.60Qwen3\.5\-2B0\.472420\.950\.493417\.820\.414312\.90Qwen3\.5\-4B0\.602130\.580\.612227\.870\.510318\.59Qwen3\.5\-9B0\.668836\.730\.680134\.960\.556622\.19Qwen3\.5\-27B0\.714441\.760\.726441\.220\.585924\.93Qwen3\.5\-122B\-A10B0\.737344\.420\.749444\.320\.604526\.68Dedicated translation modelsAfriNLLB\-600M0\.731547\.670\.767149\.810\.600828\.38NLLB\-600M0\.734046\.810\.765449\.030\.602827\.68NLLB\-1\.3B0\.747148\.720\.775350\.770\.610128\.52NLLB\-3\.3B0\.752549\.790\.779552\.360\.613929\.03Hunyuan\-MT\-7B0\.505623\.490\.521119\.550\.439115\.48TranslateGemma\-4B0\.548724\.780\.569923\.120\.464415\.21TranslateGemma\-12B0\.712539\.480\.721138\.920\.598324\.55TranslateGemma\-27B0\.733543\.320\.749943\.700\.610226\.24African language\-specialized modelsAfriqueLlama\-8B0\.668937\.590\.717941\.580\.563023\.03AfriqueGemma\-4B0\.691640\.740\.722342\.490\.568823\.60AfriqueGemma\-12B0\.713144\.520\.742946\.260\.584725\.94AfriqueQwen\-8B0\.675438\.230\.716841\.080\.565923\.06AfriqueQwen\-14B0\.695541\.580\.729743\.690\.575424\.64OursTranslatePsy\-AfriSLM\-0\.8B0\.743248\.300\.777051\.300\.612028\.54TranslatePsy\-AfriSLM\-2B0\.755250\.000\.787053\.130\.618729\.32TranslatePsy\-AfriSLM\-4B0\.762951\.130\.794154\.330\.623729\.97Table 18:COMET\-22 and ChrF\+\+results on Flores\-200, BOUQuET and Smol\.Figure 13:Multi\-turn, conversational, cross\-lingual translation and identification \(TranslatePsy\-AfriSLM\-4B\)\.

Similar Articles

Sample-Size Scaling of the African Languages NLI Evaluation

arXiv cs.CL

This paper examines the effect of labeled data size on natural language inference performance for 16 African languages using the AfriXNLI benchmark. The results show that scaling behavior is language-sensitive and often non-monotonic, challenging the common assumption of monotonic improvement, and emphasizing the need for language-specific dataset creation and stronger multilingual strategies.