Score
Design and implement measurement systems that decompose culture into measurable dimensions (for example, cultural‑faithfulness axes) and derive quantitative metrics for cultural variation. Use those metrics to analyze datasets and evaluate or compare models across cultural groups, creating operationalized evaluation procedures and statistical comparisons.
This study addresses the frequent neglect of local cultural perspectives in existing automated evaluations of AI-generated images, particularly regarding “cultural appropriateness.” It introduces a novel evaluation framework that deeply integrates diverse community participation from the outset, collaborating with blind and visually impaired individuals in the UK and residents of Kerala and Tamil Nadu in India to systematically translate lived cultural experiences and community concerns into actionable assessment dimensions. Leveraging multimodal large language models as judges (LLM-as-a-judge), the approach operationalizes community consensus into structured scoring rules, enabling automated evaluation of cultural appropriateness. The work not only establishes a conceptual framework grounded in community values and demonstrates its feasibility but also exposes critical limitations in current AI models’ understanding of cultural context.
This study challenges the prevailing view of language models as passive recorders of cultural phenomena, arguing instead that they function as active apparatuses that co-constitute cultural reality. Drawing on Karen Barad’s concept of “agential cuts” and adopting a material-discursive perspective, the research integrates natural language processing, qualitative analysis, and apparatus critique—illustrated through case studies such as film and television dialogue—to uncover the entanglements inherent in how models delineate cultural boundaries. Emphasizing ethical and theoretical reflexivity in methodological choices for cultural measurement, the work proposes a new paradigm that is both culturally sensitive and theoretically grounded. It further demonstrates how current model designs often erase cultural markers and diminish historical sensitivity, thereby shaping—and potentially distorting—the production and interpretation of cultural structures.
This study addresses the current lack of systematic approaches for evaluating artificial intelligence’s adaptability and comprehension across diverse cultural contexts. Drawing on measurement theory, it introduces—for the first time—the validity framework from psychometrics into the assessment of AI cultural competence, thereby disentangling the construct of “cultural intelligence” from its operationalization. The work proposes a modular and extensible evaluation paradigm that integrates cultural dimension modeling, indicator design, data collection, and assessment protocols. By delineating core competency domains and their corresponding measurable indicators, this research establishes a theoretical and methodological foundation for large-scale, systematic evaluation of AI systems’ cultural adaptability.
This study investigates the generalization capability and fairness of large language models (LLMs) across culturally diverse measurement systems (e.g., currency, units). Addressing three core research questions—(RQ1) inherent measurement system preferences, (RQ2) accuracy disparities across systems, and (RQ3) whether reasoning mitigates bias against non-dominant systems—the authors construct a benchmark dataset covering seven open-source LLMs and diverse cross-cultural measurement scenarios. They systematically evaluate chain-of-thought (CoT) and other reasoning-based prompting techniques. The work首次 identifies “measurement system representation bias” as a source of latent computational inequity. Empirical results reveal significant dominant-system preference across all models, with substantially lower accuracy on non-dominant systems. CoT improves accuracy by up to 37%, yet increases average response length by 2.1×, exposing a cultural dimension of reasoning cost discrimination.
This study investigates whether large language models (LLMs) can dynamically adapt their value orientation and response content according to users’ national cultural values, operationalized via Hofstede’s five-dimensional cultural framework. Method: We construct culture-specific personas for 36 countries and employ multilingual prompts, complemented by cross-lingual consistency analysis and mixed qualitative–quantitative evaluation to systematically assess LLMs’ cultural recognition capability and value alignment fidelity. Results: While LLMs reliably detect surface-level cultural distinctions—e.g., the individualism–collectivism spectrum—they consistently fail to achieve deep cultural adaptation, exhibiting a “recognition–execution gap” in value alignment. To address this, we propose CultAlign, the first culturally sensitive training framework, and CultAlign-Bench, a reusable, multilingual benchmark for cross-cultural alignment evaluation. These contributions provide both methodological foundations and empirical evidence for modeling, assessing, and improving LLMs’ cultural self-adaptivity.
本文通过分解多个大型语言模型的响应变异,采用噪声信号比方法评估模型文化定位的可靠性,发现模型的文化坐标易受提问方式影响,难以精确解读。
This work addresses the limitation of existing cultural benchmarks, which predominantly assess factual knowledge while neglecting deeper reasoning capabilities such as explaining, substantiating, and revising cultural references. Using literary interpretation as the evaluation context, this study proposes an evidence-centered benchmark that systematically examines models’ deep cultural understanding through the cross-contextual identification and reconstruction of cultural references. Methodologically, it integrates literary data analysis, contextual resources, and expert feedback mechanisms. The framework is validated using Danish literature case studies to evaluate models’ cultural robustness and interpretive depth while preserving legitimate scholarly disagreement. Ultimately, this research establishes an evaluation paradigm that transcends conventional metrics, advancing AI development toward systems capable of sophisticated cultural reasoning.
This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.
This study addresses the widespread neglect of cultural fidelity in current video generation models, which predominantly prioritize visual quality. To systematically evaluate cultural accuracy, the authors propose CultureScore—the first quantifiable, multidimensional assessment framework that measures model performance across three key dimensions: identity, context, and behavior. Leveraging a dataset of 6,180 generated videos spanning culturally diverse scenarios from ten countries, and combining human annotations with automated metrics, the analysis reveals that even state-of-the-art models achieve an overall CultureScore of only 56.8%, with behavioral representation scoring lowest (<52%). Notably, human preferences align closely with CultureScore but diverge significantly from visual quality ratings, exposing a systematic deficiency in how existing models represent cultural nuances.
本文探讨了可视化素养测量框架的起源,区分了基于评估和个人熟练度的第一波与关注情境实践的第二波,并提出未来研究方向。