Tag
This large-scale study investigates GRPO in non-English and multilingual contexts, finding that training to reason in native languages has a small gap to English training and reveals strong crosslingual transfer, though effects are model- and language-dependent.
This paper investigates the computational cost of reasoning in non-English languages, using Japanese as a case study, to highlight inefficiencies in current AI systems.
GitHub released the Multilingual Repositories Dataset, covering over 40 million repositories and 80 million classification rows, with insights into non-English READMEs, issues, and PRs. Korean leads in issue text while Portuguese tops READMEs.
This paper investigates whether compact, task-specific bi-encoders fine-tuned on synthetic data from large language models can outperform general-purpose embeddings for clinical code retrieval in non-English languages, achieving state-of-the-art results on Spanish benchmarks CodiESP and DISTEMIST.