Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network

arXiv cs.CL Papers

Summary

This paper reports on a collaboration between the European Commission's Directorate-General for Translation and the European Master's in Translation network to localize the MMLU dataset into 11 European languages, creating a more inclusive benchmark for LLM evaluation while providing authentic training for translation students.

arXiv:2607.18432v1 Announce Type: new Abstract: This paper reports on a collaboration between the Directorate-General for Translation (DGT) and the European Master's in Translation (EMT) to localise the MMLU dataset into 11 European languages. Beyond creating a more inclusive benchmark for LLM evaluation, the project offers master's students authentic, project-based professional training in translation, revision, project management, and multilingual coordination, while highlighting key methodological, administrative, and workflow challenges.
Original Article
View Cached Full Text

Cached at: 07/22/26, 08:23 AM

# Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network
Source: [https://arxiv.org/abs/2607.18432](https://arxiv.org/abs/2607.18432)
[View PDF](https://arxiv.org/pdf/2607.18432)

> Abstract:This paper reports on a collaboration between the Directorate\-General for Translation \(DGT\) and the European Master's in Translation \(EMT\) to localise the MMLU dataset into 11 European languages\. Beyond creating a more inclusive benchmark for LLM evaluation, the project offers master's students authentic, project\-based professional training in translation, revision, project management, and multilingual coordination, while highlighting key methodological, administrative, and workflow challenges\.

## Submission history

From: Susana Valdez \[[view email](https://arxiv.org/show-email/62c16c4d/2607.18432)\] **\[v1\]**Mon, 20 Jul 2026 18:33:59 UTC \(456 KB\)

Similar Articles

UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding

arXiv cs.CL

UrduMMLU is a new benchmark of 26,431 multiple-choice questions across 26 subjects for evaluating LLMs on Urdu language understanding, sourced from native educational materials. Evaluation of 30 LLMs reveals Gemini-3.5-Flash performs best, while open-source models and region-specific subjects pose significant challenges.