Cross-Lingual Transfer for Machine Translation in Turkic Languages

arXiv cs.CL Papers

Summary

This paper studies cross-lingual transfer for machine translation among five Turkic languages using pairwise transfer matrices with mT5, finding that transfer is strongest between closely related pairs and that Latinization helps in script-mismatched settings.

arXiv:2607.29355v1 Announce Type: new Abstract: Cross-lingual transfer is central to low-resource machine translation, but its behavior within closely related language families remains insufficiently characterized. We study transfer among five Turkic languages; Turkish, Azerbaijani, Uzbek, Kazakh, and Kyrgyz; using pairwise transfer matrices. In this setting, each model is fine-tuned with one transfer source and evaluated on a different transfer target while the translation target remains the same. Across mT5 experiments, we find that transfer is strongest between closely related Turkic pairs, especially Turkish-Azerbaijani and Kazakh-Kyrgyz. We also show that transfer direction matters, and that the same transfer source-transfer target pair can behave differently when the translation target changes. Latinization improves BLEU and chrF in several script-mismatched settings, but its effect is not uniform across metrics. Additional analyses show that transfer sources are mostly stable across different datasets and model settings.
Original Article
View Cached Full Text

Cached at: 08/03/26, 07:36 AM

# Cross-Lingual Transfer for Machine Translation in Turkic Languages
Source: [https://arxiv.org/abs/2607.29355](https://arxiv.org/abs/2607.29355)
[View PDF](https://arxiv.org/pdf/2607.29355)

> Abstract:Cross\-lingual transfer is central to low\-resource machine translation, but its behavior within closely related language families remains insufficiently characterized\. We study transfer among five Turkic languages; Turkish, Azerbaijani, Uzbek, Kazakh, and Kyrgyz; using pairwise transfer matrices\. In this setting, each model is fine\-tuned with one transfer source and evaluated on a different transfer target while the translation target remains the same\. Across mT5 experiments, we find that transfer is strongest between closely related Turkic pairs, especially Turkish\-Azerbaijani and Kazakh\-Kyrgyz\. We also show that transfer direction matters, and that the same transfer source\-transfer target pair can behave differently when the translation target changes\. Latinization improves BLEU and chrF in several script\-mismatched settings, but its effect is not uniform across metrics\. Additional analyses show that transfer sources are mostly stable across different datasets and model settings\.

## Submission history

From: Cagri Toraman \[[view email](https://arxiv.org/show-email/db19ed5f/2607.29355)\] **\[v1\]**Fri, 31 Jul 2026 12:42:17 UTC \(650 KB\)

Similar Articles

Tokenizing Crosslingual Homographs

arXiv cs.CL

This paper investigates how multilingual tokenizers handle cross-lingual homographs (identical surface forms with different meanings across languages) and proposes a lightweight language-cue intervention that introduces language-specific characters to reduce token sharing. Experiments show modest improvements in machine translation, particularly with BPE tokenization.

Multilingual Unlearning in LLMs: Transfer, Dynamics, and Reversibility

arXiv cs.CL

This paper studies multilingual unlearning in LLMs by extending the TOFU benchmark to five languages. It finds that unlearning transfer varies by script and family, operates primarily in later decoding layers, and that a single steering direction can recover much of the suppressed knowledge across languages.