Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network
Summary
This paper reports on a collaboration between the European Commission's Directorate-General for Translation and the European Master's in Translation network to localize the MMLU dataset into 11 European languages, creating a more inclusive benchmark for LLM evaluation while providing authentic training for translation students.
View Cached Full Text
Cached at: 07/22/26, 08:23 AM
# Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network Source: [https://arxiv.org/abs/2607.18432](https://arxiv.org/abs/2607.18432) [View PDF](https://arxiv.org/pdf/2607.18432) > Abstract:This paper reports on a collaboration between the Directorate\-General for Translation \(DGT\) and the European Master's in Translation \(EMT\) to localise the MMLU dataset into 11 European languages\. Beyond creating a more inclusive benchmark for LLM evaluation, the project offers master's students authentic, project\-based professional training in translation, revision, project management, and multilingual coordination, while highlighting key methodological, administrative, and workflow challenges\. ## Submission history From: Susana Valdez \[[view email](https://arxiv.org/show-email/62c16c4d/2607.18432)\] **\[v1\]**Mon, 20 Jul 2026 18:33:59 UTC \(456 KB\)
Similar Articles
Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages
This tutorial paper provides an overview of building multilingual and multimodal LLMs for low-resource languages, covering data creation, model alignment, fine-tuning, and evaluation, with a focus on practical recipes and hands-on resources.
Multilingual Sentence Embeddings for Linguistic-Integrated Reliability Audit
This paper evaluates whether multilingual sentence embeddings can replace translation for linguistic-integrated reliability auditing across multiple languages in educational assessments, finding that native-language embeddings reproduce translation-based reliability estimates closely.
MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages
This paper introduces MultiSynt/MT, a trillion-token multilingual parallel corpus created by translating English pre-training data into 36 languages. Experiments show that LLMs trained on this translated data achieve performance comparable to native data with fewer tokens, though some evaluation blind spots and cultural gaps remain.
UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding
UrduMMLU is a new benchmark of 26,431 multiple-choice questions across 26 subjects for evaluating LLMs on Urdu language understanding, sourced from native educational materials. Evaluation of 30 LLMs reveals Gemini-3.5-Flash performs best, while open-source models and region-specific subjects pose significant challenges.
BIASEDTALES-ML: A Multilingual Dataset for Analyzing Narrative Attribute Distributions in LLM-Generated Stories
Researchers introduce BIASEDTALES-ML, a large-scale multilingual dataset of ~350,000 LLM-generated children's stories across eight languages, designed to analyze narrative attribute distributions and cross-lingual bias patterns in language model outputs. The work reveals significant cross-lingual variability, highlighting limitations of English-centric bias evaluations.