CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs

📅 2025-11-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing cultural competence evaluations predominantly rely on decontextualized correctness judgments, failing to capture the depth of understanding and reasoning that large language models (LLMs) exhibit in authentic, multicultural contexts. To address this, we propose a “thick culture” evaluation paradigm—a context-aware, scenario-based framework for cultural assessment that emphasizes situationally grounded response generation. We introduce four fine-grained, low-variance metrics—coverage, specificity, semantic depth, and coherence—to quantify cultural reasoning rigorously. Integrating contextualized benchmarks with multidimensional automated evaluation, our experiments reveal that conventional “thin” evaluation significantly overestimates model capabilities and yields high result variance. In contrast, our framework stably discriminates subtle differences in cultural understanding across state-of-the-art models, delivering more reliable, interpretable, and actionable evaluation signals.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsCognitive Modeling & Cognitive Systems: Conceptual Inference and Reasoning

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Large language models (LLMs) are increasingly deployed in culturally diverse environments, yet existing evaluations of cultural competence remain limited. Existing methods focus on de-contextualized correctness or forced-choice judgments, overlooking the need for cultural understanding and reasoning required for appropriate responses. To address this gap, we introduce a set of benchmarks that, instead of directly probing abstract norms or isolated statements, present models with realistic situational contexts that require culturally grounded reasoning. In addition to the standard Exact Match metric, we introduce four complementary metrics (Coverage, Specificity, Connotation, and Coherence) to capture different dimensions of model's response quality. Empirical analysis across frontier models reveals that thin evaluation systematically overestimates cultural competence and produces unstable assessments with high variance. In contrast, thick evaluation exposes differences in reasoning depth, reduces variance, and provides more stable, interpretable signals of cultural understanding.
Problem

Research questions and friction points this paper is trying to address.

Evaluating cultural competence in LLMs deployed across diverse environments
Addressing limitations of de-contextualized cultural correctness assessments
Developing thick evaluation methods for cultural understanding and reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Benchmarks with realistic situational contexts
Four complementary metrics for response quality
Thick evaluation for stable cultural understanding
💼 Related Jobs
No related jobs found.