Score
Designs and implements adaptations of materials, procedures, and measurement instruments so they function validly in a particular cultural or linguistic context; this includes translating and localizing content, modifying protocols to respect local norms, creating context-appropriate consent and recruitment approaches, and piloting and validating instruments with representative local participants.
Cognitive assessment tools lack standardized, statistically validated cross-cultural adaptation methodologies. Method: This study systematically evaluated adaptation practices across six multicenter studies in Europe, Asia, Africa, and South America, proposing an integrative framework combining community engagement, standardized translation protocols, and multidimensional statistical validation—including variance decomposition, diagnostic accuracy, and inter-rater reliability. Contribution/Results: The study first quantified education level (26.76%) and sociocultural–linguistic factors (6.89%) as primary sources of score variance in the MoCA-H. The Brazil-specific MMSE/BCSB adaptation achieved 94.4% sensitivity and 99.2% specificity; the Manchester Cognitive Assessment demonstrated 78.5% inter-rater agreement. Collectively, findings establish the first evidence-based, generalizable framework for culturally adapted cognitive assessment instruments.
This study addresses the systematic biases introduced by large language models in cross-cultural political discourse analysis, stemming from Anglocentrism, insufficient multilingual coverage, and narrow assumptions about political institutions—biases that undermine democratic accountability. It offers the first systematic characterization of cultural failure modes in political NLP and proposes a formal framework for cultural adaptation structured across three layers: translation, discourse, and ontology. The work further introduces an evaluation matrix grounded in cultural fidelity, calibration, and democratic safety. Through participatory data development, culture-aware transfer learning, and cross-cultural pragmatic analysis, the project establishes an actionable methodology and benchmarking system for culturally adaptive political AI, thereby providing both theoretical foundations and governance boundaries for the trustworthy and legitimate deployment of cross-cultural political artificial intelligence.
Cross-cultural questionnaire design in Information and Communication Technologies for Development (ICTD) is typically costly and time-intensive due to reliance on expert review and small-scale pilot testing. Method: This paper pioneers a systematic investigation of large language models (LLMs) for automating cross-cultural questionnaire pretesting. Starting from the U.S. Climate Opinion Survey, we employed LLM-driven text localization and cultural adaptation to generate a South Africa–contextualized version, which was evaluated alongside the direct translation in a controlled, dual-version experiment (N=116) on Prolific. Contribution/Results: The LLM-adapted version significantly outperformed the literal translation in comprehensibility and acceptability. Our work challenges the conventional expert- and pilot-dependent paradigm, empirically validating the feasibility and initial efficacy of LLM-assisted cross-cultural questionnaire design. It offers a scalable, low-cost, and efficient pathway for cultural adaptation in ICTD research.
This study investigates the capacity of large language models (LLMs) to achieve cultural adaptation in English-to-Japanese business email translation—beyond literal accuracy—to align with target-context sociopragmatic norms. We propose a culture-aware prompting framework, systematically comparing baseline translation prompts against audience-oriented, norm-guided prompts. Employing a mixed-methods approach, we combine linguistic analysis of culture-specific pragmatic patterns (e.g., honorifics, indirectness, sentence-final particles) with native Japanese speakers’ expert evaluations of register appropriateness and interpersonal tone. Results demonstrate that culturally customized prompts significantly enhance the cultural appropriateness and acceptability of LLM-generated translations in Japanese professional settings. This work constitutes the first systematic empirical validation of prompt engineering’s efficacy in shaping LLMs’ cross-cultural communicative competence. It provides both methodological guidance and empirical evidence for developing culturally sensitive, inclusive multilingual AI systems.
Current large language models (LLMs) exhibit significant deficiencies in cross-cultural social adaptability, particularly in accurately assessing social acceptability across diverse cultural contexts—whether guided by abstract values or concrete situational cues. Method: This paper introduces NormAd, the first systematic evaluation framework for quantifying LLMs’ cultural adaptability across multi-granular cultural layers—from universal values to country-specific norms. It establishes NormAd-Eti, a cross-cultural etiquette benchmark comprising 2,600 situational prompts spanning 75 countries, and employs scenario-based reasoning evaluation, multi-level prompting strategies, and human baseline comparisons. Results: Experiments reveal that even the strongest models achieve <82% accuracy under explicit normative guidance (vs. >95% for humans), plummeting to <60% when provided only with abstract values and country identifiers (vs. >90% for humans). Notably, models exhibit pronounced bias against Global South cultures, underscoring critical gaps in culturally grounded reasoning.
This study challenges the prevailing view of language models as passive recorders of cultural phenomena, arguing instead that they function as active apparatuses that co-constitute cultural reality. Drawing on Karen Barad’s concept of “agential cuts” and adopting a material-discursive perspective, the research integrates natural language processing, qualitative analysis, and apparatus critique—illustrated through case studies such as film and television dialogue—to uncover the entanglements inherent in how models delineate cultural boundaries. Emphasizing ethical and theoretical reflexivity in methodological choices for cultural measurement, the work proposes a new paradigm that is both culturally sensitive and theoretically grounded. It further demonstrates how current model designs often erase cultural markers and diminish historical sensitivity, thereby shaping—and potentially distorting—the production and interpretation of cultural structures.
论文探讨了多语言大模型在回答医学问题时应保持一致性还是适应文化差异的问题,通过调查不同国家和领域的专业人士,揭示了当前缺乏实证研究支持哪种方法更优。
研究通过实验探讨了AI生成的邮件草稿如何影响日本和美国参与者的职业邮件沟通风格,并发现AI草稿可能导致文化沟通规范被覆盖,除非它们适应用户的沟通风格。
While large language models (LLMs) can infer users’ cultural backgrounds, they struggle to proactively generate culturally adapted responses. To address this gap, this work introduces CAPRI, a multi-level dialogue dataset enriched with cultural cues, and proposes a novel evaluation framework grounded in LLMs to assess cultural reasoning capabilities alongside a quantitative cultural sensitivity metric. Systematic experiments reveal—for the first time—that without explicit step-by-step prompting, models fail to effectively apply inferred cultural knowledge (e.g., in units of measurement, temporal expressions, and numerical conventions) during response generation and exhibit a prior bias toward their country-of-origin culture. However, their adaptability improves progressively as cultural cues accumulate. CAPRI establishes a new benchmark for research on cultural adaptation in language models.
This study addresses the challenge of evaluating the contextual adaptability of large language models (LLMs) in applying values within specialized scenarios by proposing a cross-domain value alignment framework. Methodologically, it introduces a novel "value × domain" matrix derived from regulatory documents, integrating multi-domain scenario simulations, factorial experiments, and action-rationale dual evaluation to systematically assess LLMs' context-sensitive execution of honesty, autonomy, and confidentiality principles in medical and legal settings. Results indicate that while the models achieve an average behavioral appropriateness of 95.6%, their accuracy in citing domain-specific justifications fluctuates significantly. This reveals a localized failure mechanism driven by risk sensitivity, thereby underscoring the necessity for fine-grained evaluations of contextual alignment in LLMs.