Score
Designing multilingual benchmarks, experimental protocols, and localized materials (including audio and questionnaires) and accounting for language, encoding, and cultural differences when evaluating model behavior and compliance across languages.
This study addresses core challenges in multilingual benchmarking: English overrepresentation, data monopolization by high-resource countries, reliance on translation rather than localization, and misalignment with human judgment. We systematically analyze over 2,000 non-English benchmarks from 148 countries (2021–2024) via large-scale cross-lingual meta-analysis, Pearson correlation testing, and bias quantification across geographic and resource dimensions. Our first empirical finding shows that localized benchmarks exhibit significantly higher alignment with human evaluation (r = 0.68) than translated ones (r = 0.47). STEM-oriented tasks demonstrate strong correlation with human judgment (r = 0.70–0.85), whereas traditional tasks like XQuAD show weak correlation (r = 0.11–0.30). We propose a six-dimensional diagnostic framework for benchmark deficiencies, five key research directions, and three guiding principles—advancing a new paradigm of culturally adaptive, human-aligned multilingual evaluation.
This study addresses the pervasive lack of cultural competence in contemporary multilingual NLP models, which often fail to accurately interpret expressions deeply rooted in specific cultural contexts despite broad language coverage. Synthesizing insights from over 50 studies published between 2020 and 2026, the work advocates a paradigm shift from isolated language processing toward modeling the “communicative ecology,” integrating institutional norms, cultural scripts, and community practices as essential contextual dimensions. Through culturally aware evaluation benchmarks (e.g., Global-MMLU, CulturalBench), multimodal grounding of local knowledge, community-coconstructed datasets, and cultural alignment techniques, the research demonstrates that insufficient training data coverage is not the sole bottleneck—language choice, tokenization strategies, and translation benchmark design are equally critical. The paper calls for layered cultural evaluation frameworks and participatory alignment approaches to advance fair, inclusive, and culturally grounded NLP systems.
This work addresses the limitations of traditional static benchmarks—prone to saturation, contamination, and high updating costs—and the susceptibility of existing large language model (LLM) auto-scoring methods to prompt sensitivity and bias. It proposes the first three-stage framework that evaluates LLMs’ *benchmark design capability* rather than merely their question-answering performance. The approach leverages structured domain cards for extraction, quota-based multi-model collaborative item generation, and scoring via precise, numerical, and symbolic verifiers combined with psychometric analysis. From nine domains, it generates 16.7K items (retaining 15K core items) and constructs a designer–responder matrix with 152K scoring records. Empirical results reveal only a moderate correlation between design and answering abilities (Spearman ρ ≈ 0.37) and a strong negative association between invalid items and discrimination (r ≈ −0.62), demonstrating the framework’s effectiveness for scalable, cross-modal, and multilingual benchmark auditing.
Current evaluations of large language models predominantly emphasize task performance while overlooking cultural appropriateness and assessment reliability. This work proposes the first benchmark framework that integrates multicultural contexts, dynamic social interactions, and a dual-dimensional evaluation encompassing both task completion and norm adherence. The framework constructs a virtual town through location-graph-based multi-agent simulation, embedding language models as resident agents, and introduces an LLM-driven norm adjudicator alongside a mechanism for quantifying evaluator uncertainty. Experimental results reveal cross-cultural robustness disparities among models, trade-offs between task success and norm compliance, and delineate the effective boundaries of automated evaluation while underscoring the necessity of human oversight.
Contemporary large multimodal models (LMMs) suffer from narrow cultural coverage, weak support for low-resource languages, and insufficient cross-cultural visual–linguistic reasoning capabilities. To address these limitations, we introduce ALM-bench—the first multimodal evaluation benchmark covering 100 languages (including numerous low-resource ones) and 13 cultural dimensions. Methodologically, we propose a hierarchical question-type design (true/false, multiple-choice, open-ended QA), integrate human-annotated multilingual image–text pairs, employ cultural knowledge graphs to guide content sampling, and establish a standardized evaluation protocol enabling multi-granularity assessment. Comprehensive experiments on leading open- and closed-source LMMs systematically expose their significant performance deficits on low-resource language understanding and culture-specific reasoning tasks—revealing these shortcomings for the first time at scale. This work advances the development of globally accessible, culturally inclusive LMMs and provides both a novel paradigm and foundational infrastructure for cross-cultural multimodal understanding research.
This study addresses three key challenges in evaluating multilingual large language models: high annotation costs, difficulty in identifying translation errors, and conflation of culture-specific knowledge with general reasoning ability. To overcome these issues, the authors propose Multilingual-IRT, the first extension of Item Response Theory (IRT) to multilingual settings. This framework introduces language-specific difficulty offsets, decomposes discrimination parameters into content- and language-related components, and models residual language proficiency effects to enable efficient and accurate assessment. Evaluated on MMLU-Pro-X across 25 models and 29 languages, Multilingual-IRT reduces prediction error by 11–16%, successfully detects translation errors in all 28 non-English languages, and identifies culture-specific items that traditional evaluation methods overlook.
This study addresses the lack of systematic evaluation of multilingual large language models (LLMs) regarding instruction hierarchy (IH) compliance, particularly their ability to consistently follow high-priority instructions in both cross-lingual and intra-lingual conflict scenarios. The authors introduce XIH-Bench, a multilingual IH evaluation benchmark spanning six languages, four domains, and three IH configurations. Their analysis reveals a previously undocumented “language boundary effect,” wherein cross-lingual conflicts paradoxically enhance IH compliance. Furthermore, models exhibit greater difficulty overriding low-priority instructions in their preferred languages, posing potential safety risks. Extensive experiments across multiple mainstream LLMs demonstrate that language choice significantly influences IH adherence, offering critical empirical insights for improving the controllability and safety of multilingual AI systems.
This work addresses the limited fine-grained diagnostic capability of existing large-scale multilingual evaluations, which hinders effective model optimization. The authors propose the first reusable multilingual agent-based diagnostic framework, decomposing post-evaluation analysis into five stages: planning, aggregation, instance inspection, cross-cultural reflection, and report generation. Integrating an expert knowledge base with multilingual understanding and culture-aware modules, the framework enables deep attribution across 33 model families, 11 benchmarks, 26 languages, and 34 cultural contexts. Leveraging an expert-driven diagnostic set comprising 54 queries across 15 languages, the approach translates scores into actionable guidance, yielding diagnostic reports that outperform the strongest baseline by 47% in quality and prevail in 87.9% of pairwise comparisons against human experts. The study further distills four key insights regarding deployment strategies, iterative refinement, and cross-cultural risk mitigation.
This study systematically evaluates the generalization of commonsense knowledge in large language models across multilingual and multicultural contexts, with a particular focus on low-resource languages and underrepresented cultures. Building upon a human-curated extension of the BLEnD benchmark encompassing over 30 language–culture pairs, the evaluation features two tracks—short-answer and multiple-choice—and strictly enforces a zero-shot setting, prohibiting any training or fine-tuning on the benchmark data while allowing participation from any NLP system. As the first large-scale, purely evaluative benchmark for cross-cultural commonsense reasoning, the initiative attracted registrations from over 140 teams, with 62 submitting results. Analysis reveals that state-of-the-art approaches perform substantially worse on low-resource languages, highlighting critical challenges in cultural alignment and cross-cultural commonsense transfer.