Tag
This large-scale study investigates GRPO in non-English and multilingual contexts, finding that training to reason in native languages has a small gap to English training and reveals strong crosslingual transfer, though effects are model- and language-dependent.