Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of current machine translation evaluation practices, which are predominantly English-centric, overlook regional cultural differences, and are susceptible to data contamination, thereby failing to assess model robustness on localized content. To remedy this, the authors propose a source-language contrastive evaluation paradigm and introduce Cultivar—a benchmark derived from a localized subset of FLORES—that compares model performance on localized versus non-localized translations to detect data contamination and evaluate regional adaptability. This framework extends the unit of evaluation from language pairs to localized content, systematically uncovering performance disparities across regional contexts. Experiments on 32 open-source models reveal insufficient robustness in specialized translation systems, evidence of overfitting to FLORES in some cases, and a consistent bias favoring U.S.-localized content.
📝 Abstract
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.
Problem

Research questions and friction points this paper is trying to address.

data contamination
localisation robustness
multilingual translation
locale-specific evaluation
translation benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

source-contrastive evaluation
localisation robustness
data contamination
Cultivar benchmark
locale-oriented translation
🔎 Similar Papers
No similar papers found.
Pinzhen Chen
Pinzhen Chen
University of Edinburgh
large language modelsLLM post-trainingmachine translationmultilinguality
K
Koel Dutta Chowdhury
University of Technology Nuremberg
X
Xiaoya Xu
Queen’s University Belfast
D
David Tan
Saarland University
D
Doreen Osmelak
Saarland University, University of Melbourne
Ona de Gibert
Ona de Gibert
PhD Student @ University of Helsinki
Machine TranslationMultilingualityKnowledge Distillation
A
Ariun-Erdene Tumurchuluun
IIIT Hyderabad, TCS Research
Ashok Urlana
Ashok Urlana
PhD Fellow at IIIT-Hyderabad and Researcher at TCS Research
LLMs attributionExplainabilityNLPMachine learningRobotics
F
Fedor Sizov
Saarland University
H
Hale Sirin
Johns Hopkins University
J
Jesujoba Alabi
Saarland University
K
Karrar Talib Abed
Imam Ja’afar Al-Sadiq University
Mateusz Klimaszewski
Mateusz Klimaszewski
PhD Student, Warsaw University of Technology
natural language processingmachine learningmachine translation
Nikolay Bogoychev
Nikolay Bogoychev
Research Scientist at Meta
Massive parallelismGPGPULLMsMachine TranslationNatural languages
Niyati Bafna
Niyati Bafna
Johns Hopkins University, Center for Language and Speech Processing
Low-resource NLPLarge Language ModellingMachine TranslationBilingual Lexicon Induction
P
Patricia Schmidtova
Charles University
P
Preksha Manjunath Shanbhag
Queen’s University Belfast
S
Sherrie Shen
University of Edinburgh
V
Vilem Zouhar
ETH Zurich
Vivek Iyer
Vivek Iyer
Division of Cardiology, Department of Medicine, Columbia University Medical Center
Cardiac ElectrophysiologyIon Channel BiophysicsComputational BiologyClinical ElectrophysiologyCalcium Cycling
Y
Yasser Hamidullah
University of Zurich
Y
Yusser Al Ghussin
Saarland University
Zheng Zhao
Zheng Zhao
PhD Student, University of Edinburgh
natural language processinglarge language modelscomputational linguistics