Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network
Summary
This paper reports on a collaboration between the European Commission's Directorate-General for Translation and the European Master's in Translation network to localize the MMLU dataset into 11 European languages, creating a more inclusive benchmark for LLM evaluation while providing authentic training for translation students.
View Cached Full Text
Cached at: 07/22/26, 08:23 AM
# Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network Source: [https://arxiv.org/abs/2607.18432](https://arxiv.org/abs/2607.18432) [View PDF](https://arxiv.org/pdf/2607.18432) > Abstract:This paper reports on a collaboration between the Directorate\-General for Translation \(DGT\) and the European Master's in Translation \(EMT\) to localise the MMLU dataset into 11 European languages\. Beyond creating a more inclusive benchmark for LLM evaluation, the project offers master's students authentic, project\-based professional training in translation, revision, project management, and multilingual coordination, while highlighting key methodological, administrative, and workflow challenges\. ## Submission history From: Susana Valdez \[[view email](https://arxiv.org/show-email/62c16c4d/2607.18432)\] **\[v1\]**Mon, 20 Jul 2026 18:33:59 UTC \(456 KB\)
Similar Articles
Predicting Multilingual Classification and Translation Performance of LLMs with Cross-Lingual Alignment $\unicode{x2013}$ Is English Enough?
The paper compares 27 cross-lingual alignment (CLA) score variants for predicting LLM performance on multilingual classification and translation tasks, and proposes a PMI-based translation metric. It finds that CLA with English predicts translation quality comparably to or better than source-target CLA, supporting the view that LLMs use English as an internal pivot language.
Analysis of Numerical Localisation in LLM Translations
This paper analyses the capability of five large language models to localise times, numbers, and dates when translating between English and German, and tests strategies to improve accuracy—finding that embedding localisation principles into the prompt context yields statistically significant improvements.
Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages
This tutorial paper provides an overview of building multilingual and multimodal LLMs for low-resource languages, covering data creation, model alignment, fine-tuning, and evaluation, with a focus on practical recipes and hands-on resources.
Multilingual Sentence Embeddings for Linguistic-Integrated Reliability Audit
This paper evaluates whether multilingual sentence embeddings can replace translation for linguistic-integrated reliability auditing across multiple languages in educational assessments, finding that native-language embeddings reproduce translation-based reliability estimates closely.
MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages
This paper introduces MultiSynt/MT, a trillion-token multilingual parallel corpus created by translating English pre-training data into 36 languages. Experiments show that LLMs trained on this translated data achieve performance comparable to native data with fewer tokens, though some evaluation blind spots and cultural gaps remain.