Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network

arXiv cs.CL Papers

Summary

This paper reports on a collaboration between the European Commission's Directorate-General for Translation and the European Master's in Translation network to localize the MMLU dataset into 11 European languages, creating a more inclusive benchmark for LLM evaluation while providing authentic training for translation students.

arXiv:2607.18432v1 Announce Type: new Abstract: This paper reports on a collaboration between the Directorate-General for Translation (DGT) and the European Master's in Translation (EMT) to localise the MMLU dataset into 11 European languages. Beyond creating a more inclusive benchmark for LLM evaluation, the project offers master's students authentic, project-based professional training in translation, revision, project management, and multilingual coordination, while highlighting key methodological, administrative, and workflow challenges.
Original Article
View Cached Full Text

Cached at: 07/22/26, 08:23 AM

# Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network
Source: [https://arxiv.org/abs/2607.18432](https://arxiv.org/abs/2607.18432)
[View PDF](https://arxiv.org/pdf/2607.18432)

> Abstract:This paper reports on a collaboration between the Directorate\-General for Translation \(DGT\) and the European Master's in Translation \(EMT\) to localise the MMLU dataset into 11 European languages\. Beyond creating a more inclusive benchmark for LLM evaluation, the project offers master's students authentic, project\-based professional training in translation, revision, project management, and multilingual coordination, while highlighting key methodological, administrative, and workflow challenges\.

## Submission history

From: Susana Valdez \[[view email](https://arxiv.org/show-email/62c16c4d/2607.18432)\] **\[v1\]**Mon, 20 Jul 2026 18:33:59 UTC \(456 KB\)

Similar Articles

Predicting Multilingual Classification and Translation Performance of LLMs with Cross-Lingual Alignment $\unicode{x2013}$ Is English Enough?

arXiv cs.CL

The paper compares 27 cross-lingual alignment (CLA) score variants for predicting LLM performance on multilingual classification and translation tasks, and proposes a PMI-based translation metric. It finds that CLA with English predicts translation quality comparably to or better than source-target CLA, supporting the view that LLMs use English as an internal pivot language.

Analysis of Numerical Localisation in LLM Translations

arXiv cs.CL

This paper analyses the capability of five large language models to localise times, numbers, and dates when translating between English and German, and tests strategies to improve accuracy—finding that embedding localisation principles into the prompt context yields statistically significant improvements.