Crosslingual On-Policy Self-Distillation for Multilingual Reasoning
Summary
The paper proposes Crosslingual On-Policy Self-Distillation (COPSD), a method to transfer high-resource language reasoning capabilities to low-resource languages using a shared student-teacher architecture. Experiments across 17 African languages show significant improvements in mathematical reasoning and answer-format adherence, outperforming Group Relative Policy Optimization (GRPO).
View Cached Full Text
Cached at: 05/12/26, 10:51 AM
Paper page - Crosslingual On-Policy Self-Distillation for Multilingual Reasoning
Source: https://huggingface.co/papers/2605.09548
Abstract
COPSD transfers high-resource language model reasoning behavior to low-resource languages using self-distillation with crosslingual context, improving mathematical reasoning performance.
Large language models(LLMs) have achieved remarkable progress inmathematical reasoning, but this ability is not equally accessible across languages. Especiallylow-resource languagesexhibit much lower reasoning performance. To address this, we proposeCrosslingual On-Policy Self-Distillation(COPSD), which transfers a model’s own high-resource reasoning behavior tolow-resource languages. COPSD uses the same model as student and teacher: the student sees only the low-resource problem, while the teacher receives privileged crosslingual context, including the problem translation and reference solution in English. Training minimizes full-distributiontoken-level divergenceon the student’s own rollouts, providing dense supervision while avoiding the sparsity and instability of outcome-onlyreinforcement learning(RL). Experiments on 17 low-resource African languages show that COPSD consistently improves low-resourcemathematical reasoningacross model sizes and substantially outperforms Group RelativePolicy Optimization(GRPO). Further analyses show that COPSD improves answer-format adherence, strengthenstest-time scaling, and generalizes to hardermultilingual reasoning benchmarks, with especially large gains for lower-resource languages. We make our code and data available at: https://github.com/cisnlp/COPSD.
View arXiv pageView PDFGitHubAdd to collection
Get this paper in your agent:
hf papers read 2605\.09548
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.09548 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.09548 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.09548 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
On-Policy Delta Distillation for Multilingual Math Reasoning
This paper studies On-Policy Delta Distillation (OPD^2) for multilingual math reasoning in English, Korean, and Japanese, showing consistent improvements over standard OPD and narrowing language gaps.
Native Multilingual Chain-of-Thought Reasoning in Low-Resource Southeast Asian Languages
Introduces OSCD, a post-training algorithm to improve native multilingual chain-of-thought reasoning in low-resource Southeast Asian languages, achieving up to 3.2x improvements on math benchmarks.
Efficient Multilingual Reasoning Transfer via Progressive Code-Switching
This paper introduces Progressive Code-Switching (PCS), a reinforcement learning approach with curriculum learning that gradually increases code-switching in LLMs to efficiently transfer multilingual reasoning capabilities.
Reasoning Compression with Mixed-Policy Distillation
This paper proposes Mixed-Policy Distillation (MPD), a framework that transfers concise reasoning behaviors from large teacher models to smaller student models, reducing token usage by up to 27.1% while improving performance.
dOPSD: On-Policy Self-Distillation for Diffusion Language Models
This paper introduces dOPSD, an on-policy self-distillation method for diffusion language models that leverages internal denoising trajectories to improve mathematical reasoning and code generation.