Tag
This paper investigates whether the native-vs-translate multilingual reasoning gap on MGSM is an artifact of output-token budgets. The authors show that the measured gap swings significantly across different hard caps and that length normalization can reverse strategy rankings, concluding that output caps should be treated as an independent variable in evaluations.
This paper revisits the multilingual reasoning gap in LLMs, finding it smaller than previously reported under comparable supervision. It introduces Layer Swap, which transfers mid-layer weights from an English reasoning specialist to native language specialists, nearly closing the gap while preserving native-language chain-of-thought.