Cross-Lingual Transfer for Machine Translation in Turkic Languages
Summary
This paper studies cross-lingual transfer for machine translation among five Turkic languages using pairwise transfer matrices with mT5, finding that transfer is strongest between closely related pairs and that Latinization helps in script-mismatched settings.
View Cached Full Text
Cached at: 08/03/26, 07:36 AM
# Cross-Lingual Transfer for Machine Translation in Turkic Languages Source: [https://arxiv.org/abs/2607.29355](https://arxiv.org/abs/2607.29355) [View PDF](https://arxiv.org/pdf/2607.29355) > Abstract:Cross\-lingual transfer is central to low\-resource machine translation, but its behavior within closely related language families remains insufficiently characterized\. We study transfer among five Turkic languages; Turkish, Azerbaijani, Uzbek, Kazakh, and Kyrgyz; using pairwise transfer matrices\. In this setting, each model is fine\-tuned with one transfer source and evaluated on a different transfer target while the translation target remains the same\. Across mT5 experiments, we find that transfer is strongest between closely related Turkic pairs, especially Turkish\-Azerbaijani and Kazakh\-Kyrgyz\. We also show that transfer direction matters, and that the same transfer source\-transfer target pair can behave differently when the translation target changes\. Latinization improves BLEU and chrF in several script\-mismatched settings, but its effect is not uniform across metrics\. Additional analyses show that transfer sources are mostly stable across different datasets and model settings\. ## Submission history From: Cagri Toraman \[[view email](https://arxiv.org/show-email/db19ed5f/2607.29355)\] **\[v1\]**Fri, 31 Jul 2026 12:42:17 UTC \(650 KB\)
Similar Articles
Biomedical Machine Translation for Low-Resource Arabic-Script Languages via Cross-Lingual Transfer and LoRA Adapter Merging
This paper presents a systematic study of cross-lingual transfer for biomedical machine translation into low-resource Arabic-script languages, using Arabic and Persian as pivots. The authors evaluate LoRA adapter merging as a zero-data transfer strategy, showing it works surprisingly well for closely related languages like Dari.
Tokenizing Crosslingual Homographs
This paper investigates how multilingual tokenizers handle cross-lingual homographs (identical surface forms with different meanings across languages) and proposes a lightweight language-cue intervention that introduces language-specific characters to reduce token sharing. Experiments show modest improvements in machine translation, particularly with BPE tokenization.
Multilingual Unlearning in LLMs: Transfer, Dynamics, and Reversibility
This paper studies multilingual unlearning in LLMs by extending the TOFU benchmark to five languages. It finds that unlearning transfer varies by script and family, operates primarily in later decoding layers, and that a single steering direction can recover much of the suppressed knowledge across languages.
Cross Lingual Transfer in Tulu Legal Comprehension: Script-Dependent Improvement and RAG-Induced Knowledge Conflict
The study evaluates cross-lingual transfer in legal comprehension for Tulu using transliteration and RAG with Kannada legal papers, revealing script-dependent improvements and challenges like fact substitution and confabulation.
MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages
This paper introduces MultiSynt/MT, a trillion-token multilingual parallel corpus created by translating English pre-training data into 36 languages. Experiments show that LLMs trained on this translated data achieve performance comparable to native data with fewer tokens, though some evaluation blind spots and cultural gaps remain.