阿拉伯语-俄语机器翻译基准测试:微调NMT与少样本LLM在丰富形态和低词汇重叠下的对比

arXiv cs.CL 论文

摘要

本文通过对比微调的NMT模型和少样本LLM,对阿拉伯语-俄语机器翻译进行基准测试,发现在低资源条件下,微调的NMT显著优于LLM。

arXiv:2609.29559v1 Announce Type: new Abstract: Arabic-Russian machine translation (MT) remains under-explored due to the rich morphology of Arabic and low lexical overlap between the two languages. We benchmark seven fine-tuned neural machine translation (NMT) models against four few-shot large language models (LLMs) on a 20k/5k/5k split of a new 15.47M-pair corpus. Fine-tuned NLLB-1.3B achieves the highest BLEU (16.3) and COMET (0.738). Aya-Expanse 8B leads the few-shot LLMs (BLEU 1.7 on 500 sentences, chrF 25.7), but all LLM scores remain far below the fine-tuned NMT baselines. Error analysis identifies low lexical overlap as the dominant failure mode; among the worst translations, mT5-small produces 32% too-short outputs. Bootstrap tests confirm significant differences among most models. Our results demonstrate that fine-tuned NMT significantly outperforms few-shot LLMs for Arabic-Russian translation under low-resource conditions.
查看原文
查看缓存全文

缓存时间: 2026/09/25 09:24

# Benchmarking Arabic–Russian Machine Translation: A Comparison of Fine-tuned NMT and Few-shot LLMs under Rich Morphology and Low Lexical Overlap
Source: [https://arxiv.org/html/2609.29559](https://arxiv.org/html/2609.29559)
Mullosharaf K\. Arabov[https://orcid.org/0000-0003-2525-1183](https://orcid.org/0000-0003-2525-1183)††thanks:Email:MKArabov@kpfu\.ruAffiliation:Kazan Federal University, Institute of Computational Mathematics and Information Technologies, Kazan, Russia

###### Abstract

Arabic–Russian machine translation \(MT\) remains under\-explored due to the rich morphology of Arabic and low lexical overlap between the two languages\. We benchmark seven fine\-tuned neural machine translation \(NMT\) models against four few\-shot large language models \(LLMs\) on a 20k/5k/5k split of a new 15\.47M\-pair corpus\. Fine\-tuned NLLB\-1\.3B achieves the highest BLEU \(16\.3\) and COMET \(0\.738\)\. Aya\-Expanse 8B leads the few\-shot LLMs \(BLEU 1\.7 on 500 sentences, chrF 25\.7\), but all LLM scores remain far below the fine\-tuned NMT baselines\. Error analysis identifies low lexical overlap as the dominant failure mode; among the worst translations, mT5\-small produces 32% too\-short outputs\. Bootstrap tests confirm significant differences among most models\. Our results demonstrate that fine\-tuned NMT significantly outperforms few\-shot LLMs for Arabic–Russian translation under low\-resource conditions\.

Keywords:Arabic–Russian machine translation, low\-resource MT, fine\-tuning, few\-shot LLMs, NLLB, mT5, LoRA, QLoRA, COMET, BLEU, morphological complexity, lexical overlap

## 1Introduction

Arabic–Russian machine translation \(MT\) remains under\-researched despite its geopolitical and practical importance\. The recently released Arabic–Russian Translation Corpus\([Arabov, 2026c](https://arxiv.org/html/2609.29559#bib.bib1)\)provides more than 15\.8 million parallel sentence pairs, and its curated subset\([ArabicNLPWorld, 2026](https://arxiv.org/html/2609.29559#bib.bib2)\)adds about 116k high\-quality pairs from six domains\. These resources make it possible, for the first time, to systematically benchmark Arabic–Russian MT under realistic low\-resource conditions\.

Arabic is a morphologically rich language with a non\-concatenative root\-and\-pattern system, complex cliticization, and substantial dialectal variation\. These properties are described in detail by[Watson \(2007\)](https://arxiv.org/html/2609.29559#bib.bib6), who focuses on Arabic phonology and morphology, and by[Holes and Allen \(2004\)](https://arxiv.org/html/2609.29559#bib.bib7), who covers the structures, functions, and varieties of Modern Arabic\. Together with the typological distance between Arabic and Russian, these features create two central challenges for MT: rich Arabic morphology and low lexical overlap between Arabic and Russian\.

In this work, we benchmark Arabic–Russian translation using a 20k training split under realistic low\-resource conditions\. We compare seven fine\-tuned NMT models, trained with full fine\-tuning, LoRA, or QLoRA, against four instruction\-tuned LLMs evaluated in zero\-shot and few\-shot settings\. For LLMs, we use a three\-stage design: \(i\) zero\-shot on 500 sentences, \(ii\) few\-shot with five examples on the same 500 sentences to rank the models, and \(iii\) full\-scale evaluation of the best LLM on 5,000 sentences\. Our fine\-tuning strategy builds on prior work on parameter\-efficient adaptation of large language models to low\-resource languages, such as the comparative study of LoRA and QLoRA for Bashkir by[Arabov and Khaybullina \(2026\)](https://arxiv.org/html/2609.29559#bib.bib4)\.

On a 20k/5k/5k split, NLLB\-1\.3B achieves the best performance among fine\-tuned NMT models, with BLEU 16\.3 and COMET 0\.738\. Among few\-shot LLMs, Aya\-Expanse 8B leads with chrF 25\.7, BERTScore 0\.675, and COMET 0\.654, but remains substantially below the best fine\-tuned NMT system\. Full\-scale evaluation of Aya\-Expanse on 5,000 sentences confirms this gap: BLEU 5\.6 and COMET 0\.612\. Error analysis shows that low lexical overlap is the dominant failure mode, while among the worst translations mT5\-small produces 32% too\-short outputs, which we hypothesize is due to the morphological complexity of Arabic\.

Our contributions are: \(1\) a benchmark of seven fine\-tuned NMT models and four few\-shot LLMs for Arabic–Russian translation; \(2\) statistical evaluation with bootstrap confidence intervals and paired significance tests; \(3\) an error taxonomy for low\-resource Arabic–Russian MT; and \(4\) an open\-source toolkit for reproduction\.

We do not propose a new model architecture\. Instead, this work provides a statistically rigorous benchmark for Arabic–Russian translation, filling a significant gap in under\-resourced MT research\. Our study also complements previous Arabic–Russian efforts, such as the scientific\-domain parallel corpus and LLM benchmark introduced by[Arabov \(2026b\)](https://arxiv.org/html/2609.29559#bib.bib5), and relates to work on transliteration between Arabic\-script and Cyrillic\-script languages\([Arabov, 2026a](https://arxiv.org/html/2609.29559#bib.bib3)\)\.

## 2Related Work

#### Arabic linguistic complexity\.

The linguistic complexity of Arabic has been extensively documented by[Watson \(2007\)](https://arxiv.org/html/2609.29559#bib.bib6), who provides a comprehensive account of Arabic phonology and morphology, and by[Holes and Allen \(2004\)](https://arxiv.org/html/2609.29559#bib.bib7), who describes the structures, functions, and varieties of Modern Arabic\. These works highlight the main obstacles for Arabic NLP: orthographic ambiguity, morphological complexity, and the coexistence of Modern Standard Arabic with numerous dialects\.

#### Arabic preprocessing and morphological tools\.

To mitigate the challenges caused by Arabic morphology, several tools have been developed\.[Habash and Rambow \(2005\)](https://arxiv.org/html/2609.29559#bib.bib8)propose a joint model for Arabic tokenization, part\-of\-speech tagging, and morphological disambiguation\. MADAMIRA\([Pasha et al\., 2014](https://arxiv.org/html/2609.29559#bib.bib9)\)offers fast and comprehensive morphological analysis and disambiguation, while Farasa\([Darwish and Mubarak, 2016](https://arxiv.org/html/2609.29559#bib.bib10)\)provides a fast Arabic word segmenter suitable for large\-scale processing\. A recent survey by[Alrekabee \(2025\)](https://arxiv.org/html/2609.29559#bib.bib11)reviews modern Arabic preprocessing and representation techniques, including subword tokenization and contextual embeddings, and discusses their interaction with pretrained language models\.

#### Arabic corpora and parallel resources\.

The availability of Arabic corpora has been surveyed by[Zaghouani \(2014\)](https://arxiv.org/html/2609.29559#bib.bib13), who identified 66 freely available sources and highlighted the lack of comprehensive, updated resources\. Since then, several new datasets have appeared\. In this work, we use the Arabic–Russian Translation Corpus\([Arabov, 2026c](https://arxiv.org/html/2609.29559#bib.bib1)\), a large parallel resource with more than 15\.8 million pairs, and its curated subset\([ArabicNLPWorld, 2026](https://arxiv.org/html/2609.29559#bib.bib2)\), which contains about 116k pairs from religion, dictionaries, the Bible, Tatoeba, news, and conversation domains\. For dialectal Arabic, the MADAR corpus and lexicon\([Bouamor et al\., 2018](https://arxiv.org/html/2609.29559#bib.bib12)\)is a valuable resource for studying dialectal variation, which is relevant because dialectal text can appear even in predominantly MSA corpora\.

#### Previous Arabic–Russian MT work\.

A closely related study by[Arabov \(2026b\)](https://arxiv.org/html/2609.29559#bib.bib5)introduced a smaller Arabic–Russian scientific\-domain parallel corpus of about 27k pairs and benchmarked three multilingual models using LoRA/QLoRA\. The authors found that fine\-tuned models outperform few\-shot prompting and that domain\-specific fine\-tuning is necessary for scientific texts\. Our work extends this line of research by using a much larger general\-domain corpus and by systematically comparing a wider range of fine\-tuned NMT models with few\-shot LLMs\.

#### Parameter\-efficient fine\-tuning for low\-resource languages\.

LoRA and QLoRA have become standard for adapting large models to low\-resource languages\.[Arabov and Khaybullina \(2026\)](https://arxiv.org/html/2609.29559#bib.bib4)compared LoRA and QLoRA for Bashkir, a low\-resource agglutinative language, and showed that QLoRA on 7B\-scale models provides a good trade\-off between translation quality and computational cost\. We adopt the same PEFT strategies for Arabic–Russian translation\.

#### Transliteration as a related task\.

Transliteration between Arabic\-script and Cyrillic\-script languages shares with MT the challenge of handling different writing systems and phonological mismatches\.[Arabov \(2026a\)](https://arxiv.org/html/2609.29559#bib.bib3)benchmarked transliteration models for the Tajik–Farsi pair and found that byte\-level models outperform subword\-based multilingual models\. Although transliteration is distinct from translation, this work is relevant because Arabic–Russian MT must also cope with orthographic and phonological divergence between the two scripts\.

In summary, previous work has addressed Arabic linguistic preprocessing, low\-resource PEFT, Arabic corpora, and transliteration\. However, to the best of our knowledge, no prior study has systematically benchmarked Arabic–Russian MT with both fine\-tuned NMT models and few\-shot LLMs on a large general\-domain parallel corpus\. Our work fills this gap and provides a comprehensive evaluation focused on rich morphology and low lexical overlap\.

## 3Experimental Setup

Dataset\.We use the Arabic–Russian Translation Corpus\([Arabov, 2026c](https://arxiv.org/html/2609.29559#bib.bib1)\)\. The raw corpus contains 15,801,992 sentence pairs; after applying standard filtering, we retain 15,467,945 pairs for our experiments \(Table[1](https://arxiv.org/html/2609.29559#S3.T1)\)\. The majority of the filtered corpus originates from OPUS \(14,924,037\), supplemented by TED \(375,463\) and Tatoeba \(9,044\)\. We also use five manually compiled high\-quality subsets: Religion \(82,302\), Dictionary \(43,662\), Bible \(31,102\), News \(1,683\), and Conversation \(652\), all verified for alignment\.

Table 1:Composition of the filtered Arabic–Russian Translation Corpus used in our experiments\.We apply standard filtering \(empty segments, duplicates, outliers\>2,000\>2\{,\}000chars /250250words\)\. Arabic is normalised by stripping diacritics; Russian is whitespace\-normalised only\. Manual inspection of 500 random pairs confirmed 96% alignment quality\. We randomly sample 20k/5k/5k pairs for training/validation/test \(seed 42\)\. The 20k training size is constrained by computational resources and represents a realistic low\-resource scenario\.

Zero\-shot and Few\-shot LLM Prompting\.We evaluate four instruction\-tuned LLMs via Ollama:aya\-expanse:8b,llama3\.1:8b,qwen2\.5:7b\-instruct, andmistral:7b\-instruct\. All models are tested zero\-shot and few\-shot \(5 examples, fixed seed\) with the prompt:*“Output ONLY the Russian translation, nothing else\.”*Generation uses temperature 0\.1, top\-p==0\.9, and a 2048\-token context window\. For the 500\-sentence zero\-/few\-shot experiments, the maximum output length is set to 256 tokens; for the full\-scale evaluation of the best LLM on 5,000 sentences, it is increased to 512 tokens\.

LLM evaluation follows a three\-stage design: \(i\) zero\-shot on 500 sentences; \(ii\) few\-shot on the same 500 sentences to rank models; \(iii\) full\-scale evaluation of the best model on 5,000 sentences\.

Evaluation Metrics and Statistical Analysis\.We report BLEU \(sacreBLEU, case\-sensitive\), chrF \(n=6n=6\), BERTScore \(F1,bert\-base\-multilingual\-cased, Russian, rescaled\), and COMET\([Rei et al\., 2020](https://arxiv.org/html/2609.29559#bib.bib16)\)\. 95% bootstrap CIs and pairwise significance tests use 1,000 resamples \(p<0\.05p<0\.05\)\. NMT models are evaluated on the full 5,000\-sentence test set; LLMs use 500 sentences for zero\-/few\-shot comparison, with full\-scale validation on 5,000 sentences for the best\-performing LLM\.

Fine\-tuning of NMT Models\.All seven architectures are fine\-tuned on 20k sentence pairs using HuggingFace Transformers and AdamW\. Full fine\-tuning is applied to Marian\([Junczys\-Dowmunt et al\., 2018](https://arxiv.org/html/2609.29559#bib.bib17)\)and mT5\-small\([Xue et al\., 2021](https://arxiv.org/html/2609.29559#bib.bib18)\): learning rate5×10−55\{\\times\}10^\{\-5\}, batch size 16, label smoothing 0\.1\. LoRA\([Hu et al\., 2021](https://arxiv.org/html/2609.29559#bib.bib14)\)is applied to mT5\-base\([Xue et al\., 2021](https://arxiv.org/html/2609.29559#bib.bib18)\), NLLB\-600M\([Costa\-jussà et al\., 2022](https://arxiv.org/html/2609.29559#bib.bib19)\), and M2M100\([Fan et al\., 2021](https://arxiv.org/html/2609.29559#bib.bib20)\): rankr=16r\{=\}16,α=32\\alpha\{=\}32, learning rate3×10−43\{\\times\}10^\{\-4\}, batch size 8 with gradient accumulation step 2; label smoothing 0\.1 for mT5\-base, 0\.0 for NLLB\-600M and M2M100\. QLoRA\([Dettmers et al\., 2023](https://arxiv.org/html/2609.29559#bib.bib15)\)is applied to mT5\-large\([Xue et al\., 2021](https://arxiv.org/html/2609.29559#bib.bib18)\)and NLLB\-1\.3B\([Costa\-jussà et al\., 2022](https://arxiv.org/html/2609.29559#bib.bib19)\): 4\-bit quantization, rankr=32r\{=\}32,α=64\\alpha\{=\}64, learning rate2×10−42\{\\times\}10^\{\-4\}, batch size 4 with gradient accumulation step 4; label smoothing 0\.1 for mT5\-large, 0\.0 for NLLB\-1\.3B\. All models are trained for 3 epochs with 3% linear warmup; the best checkpoint is selected based on validation sacreBLEU\.

Computational Cost\.Table[2](https://arxiv.org/html/2609.29559#S3.T2)reports the wall\-clock training time and peak GPU memory for each fine\-tuned model\. All experiments were conducted on a single NVIDIA L4 GPU with 24 GB of VRAM\. As expected, full fine\-tuning of mT5\-small consumes noticeably more memory \(14\.1 GB\) than larger models adapted with LoRA/QLoRA, reflecting the overhead of updating all parameters\. QLoRA models \(mT5\-large, NLLB\-1\.3B\) require the longest training time due to 4\-bit quantisation and gradient checkpointing, but they use less memory than their full\-parameter counterparts\.

Table 2:Training time and peak GPU memory for fine\-tuned NMT models on 20k sentence pairs\.Error Classification\.We classify the 50 lowest\-scoring translations \(according to sentence\-level BLEU\) for each model into the following categories:perfect\(p=rp=r\),too\_short\(\|p\|<0\.3​\|r\|\|p\|<0\.3\|r\|\),too\_long\(\|p\|\>2\.0​\|r\|\|p\|\>2\.0\|r\|\),repetition\(most frequent token frequency\>0\.3​\|p\|\>0\.3\|p\|and at least 4 occurrences\),low\_overlap\(character\-level Jaccard similarity<0\.2<0\.2\), andother\. The thresholds are derived from the training data length distribution \(median: 22 words; 90th percentile: 48 words\) and follow prior work\([Arabov and Khaybullina, 2026](https://arxiv.org/html/2609.29559#bib.bib4)\)\. The repetition threshold is based on heuristics from MT hallucination studies\.

## 4Results

We organise the presentation of our experimental findings as follows\. First, we present the main comparison of all fine\-tuned NMT models and few\-shot LLMs on the test set, including zero\-shot baselines and a direct head\-to\-head evaluation on a common 500\-sentence subset\. Second, we report pairwise statistical significance tests to quantify the reliability of observed differences\. Third, we analyse the impact of source sentence length on translation quality\. Fourth, we provide a detailed error taxonomy for the worst translations of each model\. Fifth, we examine the correlation between different automatic metrics\. Finally, we show qualitative examples that illustrate the strengths and weaknesses of the best fine\-tuned model and the best few\-shot LLM\.

### 4\.1Main Comparison

Table[3](https://arxiv.org/html/2609.29559#S4.T3)reports BLEU, chrF, and TER for fine\-tuned NMT models on the full 5,000\-sentence test set\.NLLB\-1\.3Bachieves the highest BLEU \(16\.30±\\pm0\.50\), chrF \(34\.98±\\pm0\.61\), and lowest TER \(90\.83±\\pm0\.9\), followed by Marian and NLLB\-600M\. The mT5 family scales with size: mT5\-small fails almost completely \(BLEU 1\.51, chrF 0\.24, TER 103\.7\), mT5\-base remains weak \(BLEU 6\.54\), and mT5\-large \(BLEU 11\.28\) approaches specialised NMT performance but still lags behind M2M100 and NLLB\.

Table 3:Fine\-tuned NMT: BLEU, chrF, and TER \(mean with 95% bootstrap CI\)\. Best in each metric is bold\.Turning to semantic similarity metrics, Table[4](https://arxiv.org/html/2609.29559#S4.T4)presents BERTScore and COMET for the same models\.

Table 4:Fine\-tuned NMT: BERTScore and COMET \(mean with 95% bootstrap CI\)\. Best in each metric is bold\.NLLB\-1\.3B achieves the best COMET \(0\.738±\\pm0\.004\) and BERTScore \(0\.783±\\pm0\.002\), while Marian shows competitive performance \(COMET 0\.735±\\pm0\.004, BERTScore 0\.779±\\pm0\.002\)\. Despite its larger capacity, mT5\-large still lags behind specialised NMT architectures, with COMET 0\.682±\\pm0\.004 and BERTScore 0\.755±\\pm0\.002\.

We now turn to the evaluation of the four instruction\-tuned LLMs\. Table[5](https://arxiv.org/html/2609.29559#S4.T5)reports zero\-shot BLEU and chrF for all four models on the same 500\-sentence subset used for few\-shot evaluation\. In this setting, no single model dominates across both metrics: Aya\-Expanse 8B achieves the best BLEU \(1\.29±\\pm0\.15\), while Llama 3\.1 8B obtains the highest chrF \(21\.09±\\pm0\.47\)\. The confidence intervals for BLEU overlap among the top three models, but Aya\-Expanse and Llama 3\.1 significantly outperform Qwen 2\.5 7B and Mistral 7B\.

Table 5:Zero\-shot LLM \(500 test sentences\): BLEU and chrF with 95% bootstrap CI\. Best in each metric is bold\.The semantic metrics in Table[6](https://arxiv.org/html/2609.29559#S4.T6)complement the lexical picture\. Llama 3\.1 leads in both BERTScore \(0\.6595±\\pm0\.0023\) and COMET \(0\.6196±\\pm0\.0053\), confirming its ability to produce fluent Russian even without examples\. Aya\-Expanse does not lead in chrF, BERTScore, or COMET under zero\-shot conditions, suggesting that its advantage emerges specifically when translation examples are provided\.

Table 6:Zero\-shot LLM \(500 test sentences\): BERTScore and COMET with 95% bootstrap CI\. Best in each metric is bold\.Next, we examine the few\-shot setting\. Tables[7](https://arxiv.org/html/2609.29559#S4.T7)and[8](https://arxiv.org/html/2609.29559#S4.T8)present results for all four LLMs on the same 500 test sentences with five translation examples\. Providing five demonstrations consistently improves all metrics over the zero\-shot baseline, with Aya\-Expanse 8B becoming the clear leader across all metrics\.

Table 7:Few\-shot LLM \(5\-shot, 500 test sentences\): BLEU and chrF with 95% bootstrap CI\. Best in each metric is bold\.Aya\-Expanse achieves the best BLEU \(1\.69±\\pm0\.16\) and chrF \(25\.75±\\pm0\.65\)\. Compared with the zero\-shot condition on the same 500 sentences, Aya\-Expanse gains \+6\.31 chrF \(from 19\.44 to 25\.75\), whereas Llama 3\.1 8B improves by only \+0\.93 chrF \(from 21\.09 to 22\.02\)\. This confirms Aya\-Expanse’s specialisation for translation tasks and its ability to exploit in\-context examples more effectively than general\-purpose LLMs\. The semantic metrics in Table[8](https://arxiv.org/html/2609.29559#S4.T8)reinforce this finding\.

Table 8:Few\-shot LLM \(5\-shot, 500 sentences\): BERTScore and COMET with 95% bootstrap CI\. Best in each metric is bold\.Aya\-Expanse leads in both BERTScore \(0\.6750±\\pm0\.0030, \+0\.0293 over zero\-shot\) and COMET \(0\.6535±\\pm0\.0061, \+0\.0417\), while Llama 3\.1 shows almost no improvement in BERTScore \(\+0\.0027\) and a much smaller COMET gain \(\+0\.0151\)\. The confidence intervals for COMET do not overlap between Aya\-Expanse and Llama 3\.1 \(0\.6474–0\.6596 vs\. 0\.6289–0\.6405\), confirming a statistically significant gap\. The consistent hierarchy Aya\-Expanse \> Llama 3\.1 \> Qwen 2\.5 \> Mistral is preserved across all few\-shot metrics, with non\-overlapping CIs in most pairwise comparisons\.

### Direct NMT vs\. LLM comparison on the shared 500\-sentence subset

To enable a rigorous head\-to\-head comparison and formal significance testing, we evaluate all fine\-tuned NMT models on the same 500\-sentence subset used for the few\-shot LLM experiments\. Table[9](https://arxiv.org/html/2609.29559#S4.T9)reports BLEU and chrF scores with 95% confidence intervals, while Table[10](https://arxiv.org/html/2609.29559#S4.T10)presents BERTScore and COMET\.

Table 9:Lexical metrics \(BLEU and chrF\) on the shared 500\-sentence test subset, with 95% bootstrap CI\.On lexical metrics, the gap is extreme: NLLB\-1\.3B attains BLEU 16\.40, nearly ten times higher than Aya\-Expanse’s 1\.69\. Even the weakest NMT model, mT5\-small, yields BLEU comparable to the second\-best LLM\. The chrF scores reinforce this picture, with the best LLM reaching only 25\.75 compared to 34\.98 for NLLB\-1\.3B, and the confidence intervals do not overlap\.

Table 10:Semantic metrics \(BERTScore and COMET\) on the same 500\-sentence test subset, with 95% bootstrap CI\.The semantic metrics confirm the same pattern\. NLLB\-1\.3B achieves COMET 0\.746, compared to 0\.654 for Aya\-Expanse, with non\-overlapping confidence intervals\. Paired bootstrap tests \(1,000 resamples\) between NLLB\-1\.3B and Aya\-Expanse yieldp<0\.001p<0\.001for all four metrics \(see Section[4\.2](https://arxiv.org/html/2609.29559#S4.SS2)\)\. Even mT5\-large, which ranks fourth among the NMT systems, significantly outperforms every LLM on BLEU, chrF, and COMET \(allp<0\.001p<0\.001\)\. This direct, identically sized evaluation eliminates any confounding effect of test set size and conclusively demonstrates the superiority of fine\-tuned NMT for Arabic–Russian translation\.

To assess scalability, we evaluate Aya\-Expanse on the full 5,000\-sentence test set \(Table[11](https://arxiv.org/html/2609.29559#S4.T11)\)\. The BLEU score increases from 1\.69 to 5\.57 \(95% CI 5\.35–5\.83\), while COMET decreases from 0\.6535 to 0\.6125, confirming that the gap remains substantial: BLEU 5\.6 is roughly one\-third of NLLB\-1\.3B’s 16\.3\.

Table 11:Aya\-Expanse 8B on the full 5,000\-sentence test set \(5\-shot\)\.Overall, the main comparison demonstrates a clear hierarchy: fine\-tuned NMT models substantially outperform all few\-shot LLMs, and among LLMs, Aya\-Expanse 8B is the strongest, followed by Llama 3\.1 8B, Qwen 2\.5 7B, and Mistral 7B\. The addition of few\-shot examples improves performance across all models, but the gains are most pronounced for Aya\-Expanse, suggesting that its translation\-oriented design enables more effective use of in\-context demonstrations\.

### 4\.2Statistical Significance

We perform paired bootstrap tests \(1,000 resamples\) for each metric to determine whether the observed differences between models are statistically significant \(p<0\.05p<0\.05\)\.

Fine\-tuned NMT\.For BLEU, all pairwise differences between the top three fine\-tuned models \(NLLB\-1\.3B, Marian, NLLB\-600M\) are significant except the Marian–NLLB\-600M comparison \(p=0\.672p=0\.672\)\. The difference between NLLB\-1\.3B and Marian is significant \(p<0\.001p<0\.001\), confirming NLLB\-1\.3B as the best fine\-tuned model\. Marian and mT5\-large differ significantly \(p<0\.001p<0\.001\), as do NLLB\-600M and mT5\-large \(p<0\.001p<0\.001\)\. For chrF, Marian and NLLB\-1\.3B are not significantly different \(p=0\.126p=0\.126\), indicating similar character\-level accuracy\. NLLB\-1\.3B significantly outperforms NLLB\-600M \(p<0\.001p<0\.001\)\. For BERTScore and COMET, NLLB\-1\.3B significantly outperforms Marian \(p<0\.05p<0\.05\)\.

Table[12](https://arxiv.org/html/2609.29559#S4.T12)summarises the pairwisepp\-values for BLEU among the fine\-tuned NMT models\. Full pairwise results for all metrics are available in the supplementary material\.

Table 12:Paired bootstrap testpp\-values for BLEU among fine\-tuned NMT models \(1,000 resamples\)\.Few\-shot LLMs \(500 test sentences\)\.Paired bootstrap tests \(1,000 resamples\) for all pairwise comparisons among the four LLMs on the 500\-sentence subset show that Aya\-Expanse 8B significantly outperforms Llama 3\.1 8B across all metrics: chrF \(p<0\.01p<0\.01\), BERTScore \(p<0\.05p<0\.05\), and COMET \(p<0\.05p<0\.05\)\. Llama 3\.1 8B significantly outperforms Qwen 2\.5 7B on chrF, BERTScore, and COMET \(allp<0\.01p<0\.01\)\. Qwen 2\.5 7B significantly outperforms Mistral 7B on all three metrics \(allp<0\.001p<0\.001\)\. The full few\-shot results are reported in Tables[7](https://arxiv.org/html/2609.29559#S4.T7)and[8](https://arxiv.org/html/2609.29559#S4.T8)\. This confirms the ranking Aya \> Llama \> Qwen \> Mistral with high statistical confidence\.

NMT vs\. LLMs\.We evaluate all fine\-tuned NMT models on the same 500\-sentence test subset used for the few\-shot LLMs, enabling formal paired bootstrap tests between the two families\. The best NMT model, NLLB\-1\.3B, significantly outperforms the best few\-shot LLM, Aya\-Expanse 8B, across all four metrics: BLEU \(p<0\.001p<0\.001\), chrF \(p<0\.001p<0\.001\), BERTScore \(p<0\.001p<0\.001\), and COMET \(p<0\.001p<0\.001\)\. Even the weaker NMT models substantially outperform all LLMs: mT5\-large, which ranks fourth among the seven NMT systems, significantly exceeds every LLM on BLEU, chrF, and COMET \(allp<0\.001p<0\.001\)\.

### 4\.3Impact of Source Sentence Length

We divide the test set into five quantiles \(Q1–Q5\) by source sentence length\. All models perform best on short\-to\-medium sentences \(Q1–Q3\) and degrade on the longest 20% \(Q5\)\. The drop is more pronounced for LLMs: Aya\-Expanse loses approximately 4 BLEU points from Q3 to Q5, while NLLB\-1\.3B loses only 1\.5 points\. Few\-shot LLMs achieve their highest BLEU on sentences of 40–80 characters and drop sharply beyond 120 characters, whereas NLLB\-1\.3B remains relatively stable across all quantiles\. This suggests that explicit fine\-tuning on parallel data improves length robustness, while LLMs rely on a limited context window without length normalisation\.

### 4\.4Error Analysis

To understand the nature of translation failures, we extract the 50 worst translations \(lowest sentence BLEU\) for each model and classify them into the error types defined in Section 3 \(Experimental Setup\)\.

Table[13](https://arxiv.org/html/2609.29559#S4.T13)shows low lexical overlap as the dominant error type \(84–96%\), confirming lexical divergence as the primary bottleneck\.

Notable exceptions among fine\-tuned models are:

- •mT5\-small: 32%Shortand only 68%LowOv\. This model frequently outputs very short fragments \(e\.g\.,<extra\_id\_0\>\) due to its inability to handle long\-range dependencies\.
- •Marian: 16%Long, producing verbose translations that add unnecessary words\.
- •M2M100: 10%Long, also tending towards verbosity\.

Table 13:Error type distribution among fine\-tuned NMT models \(worst 50 translations per model, %\)\.LowOv denotes low lexical overlap \(character\-level Jaccard similarity < 0\.2\)\.Em dashes indicate 0%\.

To complement the worst\-case analysis, Table[14](https://arxiv.org/html/2609.29559#S4.T14)reports the percentage of too\-short and too\-long translations across the entire 5,000\-sentence test set\. The pattern is consistent: mT5\-small produces too\-short outputs in 48% of cases, while other models rarely fall below the length threshold\. Excessive length is infrequent but most notable in Marian \(5\.6%\) and M2M100 \(4\.3%\)\.

Table 14:Percentage of too\-short and too\-long translations across the full test set \(5,000 sentences\)\.Few\-shot LLMs exhibit almost no length or repetition errors; we did not observe any cases of token repetition in the worst\-50 analysis, consistent with the low temperature setting of 0\.1\. Nearly all errors are low lexical overlap, indicating that they generate fluent Russian but often choose synonyms or paraphrases not present in the reference\. This pattern is consistent across both the 500\-sentence few\-shot evaluation and the full 5,000\-sentence test of Aya\-Expanse\.

### 4\.5Correlation Between Metrics

We compute Spearman rank correlations between all sentence\-level metric scores aggregated over all fine\-tuned models on the full 5,000\-sentence test set\. Table[15](https://arxiv.org/html/2609.29559#S4.T15)presents the complete correlation matrix; all correlations are statistically significant atp<0\.001p<0\.001\.

Table 15:Spearman rank correlations between sentence\-level automatic metrics \(all fine\-tuned NMT models aggregated\)\.The negative correlations of TER with the other metrics are expected, as lower TER indicates better translation quality\. The strongest positive correlation is between chrF and BERTScore \(ρ=0\.855\\rho=0\.855\), followed by COMET with BERTScore \(ρ=0\.843\\rho=0\.843\) and chrF \(ρ=0\.816\\rho=0\.816\)\. BLEU shows moderate correlations with chrF \(ρ=0\.737\\rho=0\.737\) and BERTScore \(ρ=0\.763\\rho=0\.763\), indicating that it captures a somewhat different aspect of translation quality, namely exact n\-gram matches\. These relationships justify the use of multiple complementary metrics in our evaluation\.

### 4\.6Qualitative Examples

To illustrate the characteristic failure modes identified in our error analysis, Table[16](https://arxiv.org/html/2609.29559#S4.T16)presents translations of two Arabic sentences by selected models, together with the reference\. The first sentence \(*“Take, for example, ‘The Lion King’”*\) contains an idiomatic invitation followed by a proper name; the second \(*“The architects spent hundreds of hours”*\) tests lexical choice and morphological agreement\.

Table 16:Translations of two Arabic sentences by fine\-tuned NMT models and the best few\-shot LLM\.![[Uncaptioned image]](https://arxiv.org/html/2609.29559v1/tabl16.png)These examples highlight several patterns observed across our benchmark\. First,mT5\-smallcollapses completely, outputting only padding tokens or a few source words, consistent with its high rate of too\-short translations in the worst\-50 analysis\. Second,Marian and NLLB\-1\.3Bproduce lexically precise but incomplete translations: they correctly render the first clause but drop the second, sacrificing completeness for fluency\. Third,mT5\-baseintroduces spurious semantic shifts \(“thousands of hours” instead of “hundreds”\), illustrating the difficulty of controlling meaning under low\-resource conditions\. Fourth,Aya\-Expanse 8B with five examplesis the only model that captures the full meaning of the first sentence, including the idiomatic reference to*The Lion King*, but adds redundant context \(“about how I work”\) that was already present in the source\. In the second example, Aya\-Expanse’s translation is fluent and close to the reference, whereas NLLB\-1\.3B truncates the output, and Marian uses a slightly less natural verb\. These patterns confirm that low lexical overlap and morphological complexity remain the central challenges for Arabic–Russian MT, and that few\-shot LLMs trade off completeness for verbosity, while fine\-tuned NMT models err on the side of omission\.

## 5Discussion and Conclusion

Our experiments reveal a substantial and consistent performance gap between fine\-tuned NMT and few\-shot LLMs for Arabic–Russian translation\. The best fine\-tuned model, NLLB\-1\.3B, outperforms the best few\-shot LLM, Aya\-Expanse 8B, across all evaluation metrics \(BLEU 16\.3 vs\. 5\.6; COMET 0\.738 vs\. 0\.612; BERTScore 0\.783 vs\. 0\.692\)\. This difference is confirmed on the shared 500\-sentence subset, where NLLB\-1\.3B achieves COMET 0\.746 versus 0\.654 for Aya\-Expanse \(p<0\.001p<0\.001in paired bootstrap tests\)\. The consistency of this advantage across lexical and semantic metrics strongly suggests that the observed differences reflect genuine translation quality rather than metric\-specific artefacts\.

Among the LLMs, Aya\-Expanse consistently ranks first, and its lead widens in the few\-shot condition: the model gains \+6\.31 chrF and \+0\.042 COMET over its zero\-shot baseline, compared with \+0\.93 chrF and \+0\.015 COMET for Llama 3\.1\. This pattern indicates that translation\-oriented instruction tuning enables more effective use of in\-context examples than general\-purpose LLM training\. The mT5 family exhibits clear scaling behaviour: mT5\-small fails almost completely, mT5\-base performs weakly, and mT5\-large approaches but does not reach the performance of specialised NMT architectures such as NLLB and Marian\. This result is consistent with previous findings on low\-resource adaptation of large language models, where parameter\-efficient fine\-tuning of 7B\-scale models provided a favourable quality–cost trade\-off\([Arabov and Khaybullina, 2026](https://arxiv.org/html/2609.29559#bib.bib4)\)\.

The computational trade\-offs further illustrate practical considerations\. Marian, trained with full fine\-tuning, achieves a competitive BLEU of 15\.48 in only 10 minutes, whereas NLLB\-1\.3B with QLoRA requires over six hours to reach 16\.30 BLEU\. For rapid prototyping or resource\-constrained environments, Marian offers an excellent balance between quality and training cost; however, when even small improvements are critical, NLLB\-1\.3B remains preferable despite the longer training time\.

TER provides additional insight into model behaviour\. Although mT5\-base achieves a higher BLEU than mT5\-small \(6\.54 vs\. 1\.51\), its TER is comparable \(106\.7 vs\. 103\.7\), indicating that both require a similar number of edits to match the reference\. This suggests that mT5\-base produces longer but still largely incorrect translations, which BLEU masks by rewarding some n\-gram overlaps\. Thus, relying solely on BLEU may overestimate the practical usability of weak models\.

Error analysis identifies low lexical overlap as the primary failure mode, accounting for 84–96% of the worst translations among fine\-tuned NMT models\. Few\-shot LLMs produce fluent Russian but often select synonyms or paraphrases that do not match the reference, indicating a lexical\-precision deficit rather than a fluency–accuracy trade\-off\. Length robustness also differs: fine\-tuned NMT degrades gradually as source sentences become longer, whereas LLMs show a sharper drop, losing up to 4 BLEU points from the third to the fifth length quantile\. The high Spearman correlations among chrF, BERTScore, and COMET \(ρ≥0\.81\\rho\\geq 0\.81\) suggest that chrF may serve as a computationally cheaper proxy in preliminary experiments without sacrificing reliable model ranking\.

The reliability of evaluating on a 500\-sentence subset is supported by the close agreement between results on this subset and the full 5,000\-sentence test set: NLLB\-1\.3B achieves BLEU 16\.40 on the 500\-sentence subset and 16\.30 on the full set, a difference well within bootstrap confidence intervals\. This indicates that our head\-to\-head comparisons involving LLMs, which were limited to 500 sentences, provide a fair and representative estimate of relative performance\.

Although we did not conduct large\-scale human evaluation, the neural metric COMET has been shown to correlate strongly with human judgments \(r\>0\.9r\>0\.9\) in WMT shared tasks\([Rei et al\., 2020](https://arxiv.org/html/2609.29559#bib.bib16)\)\. Given that our findings are stable across multiple automatic metrics and supported by bootstrap significance testing, we consider the evidence robust\.

In summary, this study benchmarks seven fine\-tuned NMT models and four instruction\-tuned LLMs on a large Arabic–Russian parallel corpus\. The best fine\-tuned model, NLLB\-1\.3B, achieves BLEU 16\.3 and COMET 0\.738, while all LLMs remain substantially below these results even in few\-shot settings\. Low lexical overlap is the dominant challenge, and morphological complexity disproportionately affects smaller models\. Our results establish a solid baseline for the common low\-resource scenario in which only a few thousand parallel sentences are available for training, and they demonstrate that, under such conditions, dedicated NMT models remain preferable to general\-purpose LLMs for Arabic–Russian translation\.

## Limitations

Training data size\.We fine\-tuned all models on only 20k sentence pairs due to computational constraints\. Although this size is realistic for low\-resource settings, larger training sets would likely improve absolute translation quality and could alter the relative ranking of some models\.

Few\-shot evaluation scale\.Due to computational constraints, only the best\-performing LLM was evaluated on the full 5,000\-sentence test set; full\-scale evaluation of all four LLMs remains future work, although the available result for Aya\-Expanse confirms that the observed gap persists\.

Single reference translations\.All automatic metrics rely on a single Russian reference for each source sentence\. This setup can penalise valid lexical variation and may overestimate the proportion oflow\_overlaperrors, since legitimate paraphrases are not credited\.

Domain and dialect coverage\.The Arabic–Russian Translation Corpus\([Arabov, 2026c](https://arxiv.org/html/2609.29559#bib.bib1)\)is predominantly composed of Modern Standard Arabic from OPUS, with limited dialectal or domain\-specific material\. This distribution, while representative of currently available Arabic–Russian parallel data, limits the generalisability of our findings to dialectal Arabic or specialised domains\.

Error classification thresholds\.The thresholds used for error categories \(e\.g\., character\-level Jaccard similarity for low overlap, length ratios for short and long outputs\) were chosen heuristically based on the length distribution of the training data\. These thresholds may require adjustment for other language pairs or domains\.

## 6Future Work

Several directions follow from this study\. First, we plan to expand the corpus with domain\-specific dictionaries covering medicine, engineering, law, and military terminology, and to scale fine\-tuning to the full 15\.8M\-pair corpus to establish an upper\-bound performance for this language pair\. Second, explicit word\-alignment and transliteration\-based preprocessing may help reduce the lexical\-overlap bottleneck identified in our error analysis\. Third, we intend to evaluate newer multilingual models, including retrieval\-augmented approaches, and to extend the benchmark to dialectal Arabic using resources such as MADAR\([Bouamor et al\., 2018](https://arxiv.org/html/2609.29559#bib.bib12)\)\. Fourth, we plan to develop morphology\-aware evaluation metrics tailored to Arabic–Russian, building on recent surveys of Arabic preprocessing and representation\([Alrekabee, 2025](https://arxiv.org/html/2609.29559#bib.bib11)\)\. Finally, we aim to validate our findings with targeted human evaluation on a sample of low\-overlap and morphologically complex sentences\.

## References

- M\. AlrekabeeArabic NLP: a survey of pre\-processing and representation techniques\.Journal of Computer Networks, Architecture and High Performance Computing7\(4\)\.External Links:[Link](https://itscience-indexing.com/jurnal/index.php/CNAPC/article/view/6383)Cited by:[§2](https://arxiv.org/html/2609.29559#S2.SS0.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2609.29559#S6.p1.1)\.
- ArabicNLPWorld \(2026\)ArabicNLPWorldArabic\-russian parallel corpus\.Hugging Face\.Note:116,393 curated pairs, 6 sources: religion, dictionary, Bible, Tatoeba, news, conversationExternal Links:[Link](https://huggingface.co/datasets/ArabicNLPWorld/arabic-russian-parallel-corpus)Cited by:[§1](https://arxiv.org/html/2609.29559#S1.p1.1),[§2](https://arxiv.org/html/2609.29559#S2.SS0.SSS0.Px3.p1.1)\.
- Arabov and Khaybullina \(2026\)M\. K\. Arabov and S\. S\. KhaybullinaAdapting large language models to a low\-resource agglutinative language: a comparative study of LoRA and QLoRA for Bashkir\.arXiv preprintarXiv:2605\.04948\.External Links:[Link](https://arxiv.org/abs/2605.04948)Cited by:[§1](https://arxiv.org/html/2609.29559#S1.p3.1),[§2](https://arxiv.org/html/2609.29559#S2.SS0.SSS0.Px5.p1.1),[§3](https://arxiv.org/html/2609.29559#S3.p8.1),[§5](https://arxiv.org/html/2609.29559#S5.p2.1)\.
- Arabov \(2026a\)M\. K\. ArabovA systematic benchmark of machine transliteration models for the Tajik\-Farsi language pair: a comparative study from rule\-based to transformer architectures\.arXiv preprintarXiv:2605\.02270\.External Links:[Link](https://arxiv.org/abs/2605.02270)Cited by:[§1](https://arxiv.org/html/2609.29559#S1.p6.1),[§2](https://arxiv.org/html/2609.29559#S2.SS0.SSS0.Px6.p1.1)\.
- Arabov \(2026b\)M\. K\. ArabovBridging scientific heritage: an Arabic–Russian parallel corpus and LLM benchmark for sustainable knowledge transfer\.arXiv preprintarXiv:2606\.30943\.External Links:[Link](https://arxiv.org/abs/2606.30943)Cited by:[§1](https://arxiv.org/html/2609.29559#S1.p6.1),[§2](https://arxiv.org/html/2609.29559#S2.SS0.SSS0.Px4.p1.1)\.
- Arabov \(2026c\)M\. K\. ArabovArabic\-russian translation corpus\.Hugging Face\.Note:[https://huggingface\.co/datasets/ArabicNLPWorld/arabic\-russian\-translation\-corpus](https://huggingface.co/datasets/ArabicNLPWorld/arabic-russian-translation-corpus)15,801,992 sentence pairs, MIT licenseCited by:[§1](https://arxiv.org/html/2609.29559#S1.p1.1),[§2](https://arxiv.org/html/2609.29559#S2.SS0.SSS0.Px3.p1.1),[§3](https://arxiv.org/html/2609.29559#S3.p1.1),[Limitations](https://arxiv.org/html/2609.29559#Sx1.p4.1)\.
- Bouamoret al\.\(2018\)H\. Bouamor, N\. Habash, M\. Salameh, W\. Zaghouani, O\. Rambow, D\. Abdulrahim, O\. Obeid, S\. Khalifa, F\. Eryani, A\. Erdmann, and K\. OflazerThe MADAR Arabic dialect corpus and lexicon\.InProceedings of the Eleventh International Conference on Language Resources and Evaluation \(LREC 2018\),Miyazaki, Japan\.External Links:[Link](https://aclanthology.org/L18-1535/)Cited by:[§2](https://arxiv.org/html/2609.29559#S2.SS0.SSS0.Px3.p1.1),[§6](https://arxiv.org/html/2609.29559#S6.p1.1)\.
- Costa\-jussàet al\.\(2022\)M\. R\. Costa\-jussà, J\. Cross, O\. Çelebi, M\. Elbayad, K\. Heafield, K\. Heffernan, E\. Kalbassi, J\. Lam, D\. Licht, J\. Maillard, A\. Sun, S\. Wang, G\. Wenzek, A\. Youngblood, B\. Akula, L\. Barrault, G\. Mejia Gonzalez, P\. Hansanti, J\. Hoffman, S\. Jarrett, K\. R\. Sadagopan, D\. Rowe, S\. Spruit, C\. Tran, P\. Andrews, N\. F\. Ayan, S\. Bhosale, S\. Edunov, A\. Fan, C\. Gao, V\. Goswami, F\. Guzmán, P\. Koehn, A\. Mourachko, C\. Ropers, S\. Saleem, H\. Schwenk, and J\. WangNo language left behind: scaling human\-centered machine translation\.arXiv preprint\.External Links:2207\.04672,[Link](https://arxiv.org/abs/2207.04672)Cited by:[§3](https://arxiv.org/html/2609.29559#S3.p6.1)\.
- Darwish and Mubarak \(2016\)K\. Darwish and H\. MubarakFarasa: a new fast and accurate Arabic word segmenter\.InProceedings of the Tenth International Conference on Language Resources and Evaluation \(LREC’16\),Portorož, Slovenia,pp\. 1070–1074\.External Links:[Link](https://aclanthology.org/L16-1170/)Cited by:[§2](https://arxiv.org/html/2609.29559#S2.SS0.SSS0.Px2.p1.1)\.
- Dettmerset al\.\(2023\)T\. Dettmers, A\. Pagnoni, A\. Holtzman, and L\. ZettlemoyerQLORA: efficient finetuning of quantized llms\.InProceedings of the 37th International Conference on Neural Information Processing Systems,New Orleans, LA, USA\.Cited by:[§3](https://arxiv.org/html/2609.29559#S3.p6.1)\.
- Fanet al\.\(2021\)A\. Fan, S\. Bhosale, H\. Schwenk, Z\. Ma, A\. El\-Kishky, S\. Goyal, M\. Baines, O\. Celebi, G\. Wenzek, V\. Chaudhary, N\. Goyal, T\. Birch, V\. Liptchinsky, S\. Edunov, E\. Grave, M\. Auli, and A\. JoulinBeyond english\-centric multilingual machine translation\.Journal of Machine Learning Research22\(107\),pp\. 1–48\.External Links:[Link](https://jmlr.org/papers/v22/20-1307.html)Cited by:[§3](https://arxiv.org/html/2609.29559#S3.p6.1)\.
- Habash and Rambow \(2005\)N\. Habash and O\. RambowArabic tokenization, part\-of\-speech tagging and morphological disambiguation in one fell swoop\.InProceedings of the 43rd Annual Meeting of the Association for Computational Linguistics \(ACL’05\),Ann Arbor, Michigan,pp\. 573–580\.External Links:[Link](https://aclanthology.org/P05-1071/),[Document](https://dx.doi.org/10.3115/1219840.1219911)Cited by:[§2](https://arxiv.org/html/2609.29559#S2.SS0.SSS0.Px2.p1.1)\.
- Holes and Allen \(2004\)C\. Holes and R\. AllenModern arabic: structures, functions, and varieties\.Revised edition,Georgetown University Press,Washington, DC\.Cited by:[§1](https://arxiv.org/html/2609.29559#S1.p2.1),[§2](https://arxiv.org/html/2609.29559#S2.SS0.SSS0.Px1.p1.1)\.
- Huet al\.\(2021\)E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. ChenLoRA: low\-rank adaptation of large language models\.arXiv preprint\.External Links:2106\.09685,[Link](https://arxiv.org/abs/2106.09685)Cited by:[§3](https://arxiv.org/html/2609.29559#S3.p6.1)\.
- Junczys\-Dowmuntet al\.\(2018\)M\. Junczys\-Dowmunt, R\. Grundkiewicz, T\. Dwojak, H\. Hoang, K\. Heafield, T\. Neckermann, F\. Seide, U\. Germann, A\. F\. Aji, N\. Bogoychev, A\. F\. T\. Martins, and A\. BirchMarian: fast neural machine translation in C\+\+\.InProceedings of ACL 2018, System Demonstrations,Melbourne, Australia,pp\. 116–121\.External Links:[Link](https://aclanthology.org/P18-4020/),[Document](https://dx.doi.org/10.18653/v1/P18-4020)Cited by:[§3](https://arxiv.org/html/2609.29559#S3.p6.1)\.
- Pashaet al\.\(2014\)A\. Pasha, M\. Al\-Badrashiny, M\. Diab, A\. El Kholy, R\. Eskander, N\. Habash, M\. Pooleery, O\. Rambow, and R\. RothMADAMIRA: a fast, comprehensive tool for morphological analysis and disambiguation of Arabic\.InProceedings of the Ninth International Conference on Language Resources and Evaluation \(LREC’14\),Reykjavik, Iceland,pp\. 1094–1101\.External Links:[Link](https://aclanthology.org/L14-1479/)Cited by:[§2](https://arxiv.org/html/2609.29559#S2.SS0.SSS0.Px2.p1.1)\.
- Reiet al\.\(2020\)R\. Rei, C\. Stewart, A\. C\. Farinha, and A\. LavieCOMET: a neural framework for MT evaluation\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),Online,pp\. 2685–2702\.Cited by:[§3](https://arxiv.org/html/2609.29559#S3.p5.1),[§5](https://arxiv.org/html/2609.29559#S5.p7.1)\.
- Watson \(2007\)J\. C\. E\. WatsonThe phonology and morphology of arabic\.Oxford University Press,Oxford\.External Links:[Document](https://dx.doi.org/10.1093/oso/9780199257591.001.0001)Cited by:[§1](https://arxiv.org/html/2609.29559#S1.p2.1),[§2](https://arxiv.org/html/2609.29559#S2.SS0.SSS0.Px1.p1.1)\.
- Xueet al\.\(2021\)L\. Xue, N\. Constant, A\. Roberts, M\. Kale, R\. Al\-Rfou, A\. Siddhant, A\. Barua, and C\. RaffelmT5: a massively multilingual pre\-trained text\-to\-text transformer\.InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,Online,pp\. 483–498\.External Links:[Link](https://aclanthology.org/2021.naacl-main.41/),[Document](https://dx.doi.org/10.18653/v1/2021.naacl-main.41)Cited by:[§3](https://arxiv.org/html/2609.29559#S3.p6.1)\.
- Zaghouani \(2014\)W\. ZaghouaniCritical survey of the freely available arabic corpora\.arXiv preprintarXiv:1702\.07835\.External Links:[Link](https://arxiv.org/abs/1702.07835)Cited by:[§2](https://arxiv.org/html/2609.29559#S2.SS0.SSS0.Px3.p1.1)\.

相似文章

低资源Tangkhul-英语神经机器翻译

arXiv cs.CL

介绍了一个针对严重资源匮乏的Tangkhul-英语语言对的神经机器翻译系统,通过微调ByT5-large和mT5-small模型,在BLEU、chrF++、BERTScore和COMET评分上取得了优异成绩。