🤖 AI Summary
While contemporary large language models claim multilingual capabilities, the absence of evaluation benchmarks covering more than 7,000 global languages leaves over 98% of them in an assessment blind spot. This work proposes translation quality as a low-cost, scalable proxy metric for multilingual proficiency, circumventing the resource and expert dependencies inherent in traditional benchmarks. Through systematic correlation analyses across 14 prominent large language models, integrating nine multilingual downstream task benchmarks and seven translation evaluation metrics—including MetricX, xCOMET, and SSA-COMET—the study demonstrates a strong correlation between translation performance and downstream task effectiveness. For instance, the Phi-4 model exhibits median Pearson correlation coefficients ranging from 0.87 to 0.91, substantiating the validity and practicality of this approach.
📝 Abstract
The rapid proliferation of LLMs has created a critical evaluation paradox: while LLMs claim multilingual proficiency, comprehensive non-machine-translated benchmarks exist for fewer than 30 languages, leaving>98% of the world's 7,000 languages in an empirical void. Traditional benchmark construction faces scaling challenges such as cost, scarcity of domain experts, and data contamination. We evaluate the validity of a simpler alternative: can translation quality alone indicate a model's broader multilingual capabilities? Through systematic evaluation of 14 models (1B-72B parameters) across 9 diverse benchmarks and 7 translation metrics, we find that translation performance is a good indicator of downstream task success (e.g., Phi-4, median Pearson r: MetricX = 0.89, xCOMET = 0.91, SSA-COMET = 0.87). These results suggest that the representational abilities supporting faithful translation overlap with those required for multilingual understanding. Translation quality, thus emerges as a strong, inexpensive first-pass proxy of multilingual performance, enabling a translation-first screening with targeted follow-up for specific tasks.