Score
Design and execute evaluation protocols, metrics, and analyses that measure how models trained in one language or domain perform in other languages, including computing pairwise transfer scores, cross-lingual transfer matrices, and multilingual evaluation metrics to quantify accuracy gaps, asymmetries, and language-driven performance degradation. Build benchmarking suites and probing analyses to compare multilingual versus monolingual models, assess cross-lingual consistency, compare outputs to ground truth across multiple languages, and identify flip points and language-dependent failures.
This study addresses core challenges in multilingual benchmarking: English overrepresentation, data monopolization by high-resource countries, reliance on translation rather than localization, and misalignment with human judgment. We systematically analyze over 2,000 non-English benchmarks from 148 countries (2021–2024) via large-scale cross-lingual meta-analysis, Pearson correlation testing, and bias quantification across geographic and resource dimensions. Our first empirical finding shows that localized benchmarks exhibit significantly higher alignment with human evaluation (r = 0.68) than translated ones (r = 0.47). STEM-oriented tasks demonstrate strong correlation with human judgment (r = 0.70–0.85), whereas traditional tasks like XQuAD show weak correlation (r = 0.11–0.30). We propose a six-dimensional diagnostic framework for benchmark deficiencies, five key research directions, and three guiding principles—advancing a new paradigm of culturally adaptive, human-aligned multilingual evaluation.
While contemporary large language models claim multilingual capabilities, the absence of evaluation benchmarks covering more than 7,000 global languages leaves over 98% of them in an assessment blind spot. This work proposes translation quality as a low-cost, scalable proxy metric for multilingual proficiency, circumventing the resource and expert dependencies inherent in traditional benchmarks. Through systematic correlation analyses across 14 prominent large language models, integrating nine multilingual downstream task benchmarks and seven translation evaluation metrics—including MetricX, xCOMET, and SSA-COMET—the study demonstrates a strong correlation between translation performance and downstream task effectiveness. For instance, the Phi-4 model exhibits median Pearson correlation coefficients ranging from 0.87 to 0.91, substantiating the validity and practicality of this approach.
This work addresses the problem of evaluating functional consistency—i.e., whether large language models (LLMs) produce semantically equivalent outputs across languages. We propose κₚ, a novel metric grounded in response function equivalence, enabling systematic multilingual self-consistency assessment. Applying κₚ to the GlobalMMLU benchmark, we conduct the first large-scale analysis across 20 languages and 47 academic disciplines. Our results show: (1) cross-lingual self-consistency improves significantly with model parameter count; (2) a given model exhibits higher cross-lingual consistency than the inter-model consensus within the same language; and (3) κₚ effectively discriminates between model capabilities, offering an interpretable, task-agnostic benchmark for multilingual reliability. This study uncovers an intrinsic link between linguistic generalization and model scale, providing both theoretical insights and practical tools for optimizing consistency in multilingual LLM systems.
This work addresses the proliferation of large language model (LLM) evaluation benchmarks, which has outpaced systematic assessment of their intrinsic quality. To this end, we propose Benchmark², a novel framework that establishes the first quantitative methodology for evaluating the reliability and validity of LLM benchmarks through three complementary metrics: cross-benchmark ranking consistency, discriminability score, and capability alignment bias. Empirical evaluation across 15 benchmarks and 11 LLMs demonstrates that Benchmark² not only reveals substantial quality disparities among existing benchmarks but also enables the construction of streamlined test sets that maintain high evaluative performance while significantly reducing assessment scale.
Existing LLM-based evaluators exhibit language bias and unfairness in assessing non-English outputs, undermining the reliability of multilingual LLM evaluation. Method: We introduce MM-Eval, the first meta-evaluation benchmark explicitly designed for multilingual settings, covering a core set of 18 languages and a consistency set of 122 languages. It pioneers native multilingual meta-evaluation tasks—eliminating reliance on English translation—and proposes a dual-dimensional framework measuring language consistency and cross-lingual fairness, moving beyond single-metric ranking accuracy. Techniques include multilingual prompt engineering, cross-lingual consistency modeling, and absolute score distribution analysis. Contribution/Results: We empirically demonstrate that state-of-the-art English-centric evaluators suffer significant accuracy degradation and unfairness on low-resource languages. MM-Eval achieves significantly higher Best-of-N ranking correlation than existing benchmarks. All data and code are publicly released.
This study addresses underexamined translation errors in current machine translation benchmarks that may compromise the reliability and comparability of multilingual large language model (LLM) evaluations. It presents the first systematic quantification of the isolated impact of target-side translation errors on multilingual LLM assessment outcomes. The approach leverages an LLM-based evaluator to generate MQM-style error annotations, integrates the xCOMET-XXL quality estimation model, and employs controlled variable analysis while holding the correctness of English source prompts constant. Findings indicate that although automatically generated error annotations exhibit discrepancies compared to human judgments, translation errors nonetheless induce a statistically significant drop in model accuracy. This result validates the efficacy of automated error localization methods when applied to real-world translation benchmarks.
Existing evaluation methods struggle to disentangle overall performance gains in source languages from genuine cross-lingual transfer capabilities in multilingual models. To address this limitation, this work proposes the Hardness-Adjusted Transfer (HAT) score, which isolates source-language performance to more accurately quantify transfer effectiveness from high-resource to low-resource languages. Leveraging HAT, we conduct a large-scale empirical analysis across 20 language models and three major multilingual benchmarks, revealing—for the first time—that small models retain meaningful transfer capacity, that scaling model size yields diminishing returns in transfer gains, and that overall cross-lingual transfer capability has steadily improved over time.
This study addresses cross-lingual scoring bias in existing automatic machine translation evaluation metrics, a problem hindered by the lack of parallel datasets with consistent quality annotations across languages. To overcome this limitation, the authors propose XQ-MEval, the first benchmark dataset enabling cross-lingual parallel quality evaluation. Built upon the MQM error taxonomy, the dataset is constructed by automatically injecting errors into high-quality reference translations and then filtering the resulting pseudo-translations through native speakers to ensure controlled quality levels, yielding source–reference–pseudo-translation triplets. Experiments across nine language directions reveal that nine widely used metrics consistently exhibit cross-lingual biases misaligned with human judgments. The paper further introduces a score normalization strategy that substantially improves fairness and correlation with human assessments in multilingual evaluation settings.
This study addresses the challenges posed by noisy, non-parallel sentence pairs and low-quality translations prevalent in large-scale multilingual parallel corpora, as well as the absence of a unified, direction-aware evaluation framework. The authors decouple quality assessment into two distinct tasks: parallelism detection using multilingual embeddings and reference-free quality estimation employing reference-free evaluators. They further introduce a direction-aware evaluation routing mechanism to adaptively select appropriate assessment strategies. Comprehensive experiments on datasets such as FLORES-200 and BOUQuET evaluate four embedding models and nine quality estimators across diverse language directions. Results reveal that no single metric generalizes effectively across all directions, performance varies substantially by translation direction, naive ensembles dilute strong model signals, and evaluator scores correlate strongly with target-language coverage. The findings underscore the necessity of tailoring evaluation strategies to specific language directions.
Existing machine translation benchmark datasets commonly suffer from noise, structural deficiencies, and inconsistent quality, undermining the reliability of multilingual evaluation. This work proposes the first automated quality assessment framework that integrates structured corpus auditing, neural quality metrics (COMET), and fine-grained error analysis powered by large language models (LLMs) to comprehensively diagnose and refine the EU20 benchmark. By evaluating major translation systems—DeepL, Google Translate, and ChatGPT—and analyzing both reference-based and reference-free COMET scores, the study reveals a strong correlation between low COMET scores and high-accuracy errors (e.g., HellaSwag), while ARC emerges as relatively clean. The project releases a cleaned multilingual EU20 dataset, reproducible code, and a practical quality-prioritization guideline, establishing a new paradigm for constructing reliable multilingual benchmarks.