🤖 AI Summary
This work addresses the limited coverage of European languages in existing large-scale multilingual language model evaluation benchmarks, which hinders the assessment and optimization of non-English models. To bridge this gap, the project collaborates with the European Commission’s Directorate-General for Translation and the European Master’s in Translation network to localize the MMLU dataset into 11 European languages for the first time. By integrating professional translation pedagogy with a coordinated multilingual workflow—including human translation, expert review, and cross-lingual quality assurance—the effort produces a high-quality, reusable multilingual benchmark. This approach not only advances the infrastructure for multilingual model evaluation but also fosters synergies between benchmark development and translator training.
📝 Abstract
This paper reports on a collaboration between the Directorate-General for Translation (DGT) and the European Master's in Translation (EMT) to localise the MMLU dataset into 11 European languages. Beyond creating a more inclusive benchmark for LLM evaluation, the project offers master's students authentic, project-based professional training in translation, revision, project management, and multilingual coordination, while highlighting key methodological, administrative, and workflow challenges.