🤖 AI Summary
This study addresses the limitations of current machine translation evaluation practices, which are predominantly English-centric, overlook regional cultural differences, and are susceptible to data contamination, thereby failing to assess model robustness on localized content. To remedy this, the authors propose a source-language contrastive evaluation paradigm and introduce Cultivar—a benchmark derived from a localized subset of FLORES—that compares model performance on localized versus non-localized translations to detect data contamination and evaluate regional adaptability. This framework extends the unit of evaluation from language pairs to localized content, systematically uncovering performance disparities across regional contexts. Experiments on 32 open-source models reveal insufficient robustness in specialized translation systems, evidence of overfitting to FLORES in some cases, and a consistent bias favoring U.S.-localized content.
📝 Abstract
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.