Score
Adapting study materials, tasks, and protocols to local languages and cultural contexts so that semantics, usability, and ethical practices are preserved; includes designing recruitment and management procedures that produce valid, generalizable results in target populations.
Cognitive assessment tools lack standardized, statistically validated cross-cultural adaptation methodologies. Method: This study systematically evaluated adaptation practices across six multicenter studies in Europe, Asia, Africa, and South America, proposing an integrative framework combining community engagement, standardized translation protocols, and multidimensional statistical validation—including variance decomposition, diagnostic accuracy, and inter-rater reliability. Contribution/Results: The study first quantified education level (26.76%) and sociocultural–linguistic factors (6.89%) as primary sources of score variance in the MoCA-H. The Brazil-specific MMSE/BCSB adaptation achieved 94.4% sensitivity and 99.2% specificity; the Manchester Cognitive Assessment demonstrated 78.5% inter-rater agreement. Collectively, findings establish the first evidence-based, generalizable framework for culturally adapted cognitive assessment instruments.
The lack of a consensus definition of “culture” in NLP hinders systematic evaluation and cross-study comparison of culturally aware technologies. Method: We propose the first fine-grained, NLP-oriented taxonomy of cultural elements, grounded in computational linguistics, anthropology, and sociolinguistics. Through a systematic literature review and multidimensional mapping of resources, we comprehensively catalog existing culture-related datasets, models, and evaluation metrics. Contribution/Results: Our work constructs a holistic landscape spanning cultural modeling, dataset construction, and model adaptation, identifying six critical research gaps. The taxonomy unifies disparate conceptualizations of culture in NLP, providing a reusable conceptual framework and a standardized benchmarking infrastructure for culture-adaptive NLP systems.
This study addresses the systematic biases introduced by large language models in cross-cultural political discourse analysis, stemming from Anglocentrism, insufficient multilingual coverage, and narrow assumptions about political institutions—biases that undermine democratic accountability. It offers the first systematic characterization of cultural failure modes in political NLP and proposes a formal framework for cultural adaptation structured across three layers: translation, discourse, and ontology. The work further introduces an evaluation matrix grounded in cultural fidelity, calibration, and democratic safety. Through participatory data development, culture-aware transfer learning, and cross-cultural pragmatic analysis, the project establishes an actionable methodology and benchmarking system for culturally adaptive political AI, thereby providing both theoretical foundations and governance boundaries for the trustworthy and legitimate deployment of cross-cultural political artificial intelligence.
This work investigates whether regionally specialized large language models (LLMs) genuinely possess indigenous cultural understanding, using India as a case study to assess the cultural alignment of Hindi LLMs versus global models along value systems and social practices. Method: We integrate four complementary frameworks—Inglehart-Welzel Value Map, GlobalOpinionQA, CulturalBench, and NormAd—to conduct cross-model, multi-dimensional empirical evaluation. Contribution/Results: Contrary to expectations, no Hindi regional model outperforms global counterparts; some U.S. lay participants even surpass regional models on Indian value prediction tasks. Crucially, regional fine-tuning degrades—not enhances—cultural competence and harms factual knowledge recall. We identify the scarcity of high-quality, untranslated cultural data as the fundamental bottleneck. The study pioneers the argument that cultural evaluation must be co-prioritized with multilingual benchmarking and introduces an open, reusable cultural assessment methodology.
This study addresses the pervasive lack of cultural competence in contemporary multilingual NLP models, which often fail to accurately interpret expressions deeply rooted in specific cultural contexts despite broad language coverage. Synthesizing insights from over 50 studies published between 2020 and 2026, the work advocates a paradigm shift from isolated language processing toward modeling the “communicative ecology,” integrating institutional norms, cultural scripts, and community practices as essential contextual dimensions. Through culturally aware evaluation benchmarks (e.g., Global-MMLU, CulturalBench), multimodal grounding of local knowledge, community-coconstructed datasets, and cultural alignment techniques, the research demonstrates that insufficient training data coverage is not the sole bottleneck—language choice, tokenization strategies, and translation benchmark design are equally critical. The paper calls for layered cultural evaluation frameworks and participatory alignment approaches to advance fair, inclusive, and culturally grounded NLP systems.
Large language models (LLMs) often exhibit cultural misrepresentation and value conflicts due to insufficient cultural understanding—particularly pronounced in smaller models trained on limited, culturally skewed data. To address cross-cultural safety, we propose the first closed-loop framework: (1) constructing a culturally hazardous test suite and a preference alignment dataset annotated by diverse, multilingual human raters; (2) integrating cultural context modeling, multiculturally grounded human annotation, preference-informed supervised fine-tuning (SFT) and RLHF, and cross-cultural consistency evaluation. Our method enables end-to-end coverage—from quantitative cultural sensitivity assessment to fine-grained, culture-aware alignment training. Experiments demonstrate that fine-tuned small models reduce culturally insensitive outputs significantly, achieving an average 37.2% improvement in cultural sensitivity across multilingual benchmarks, while approaching the performance of large-scale models.
Current large language models (LLMs) exhibit significant deficiencies in cross-cultural social adaptability, particularly in accurately assessing social acceptability across diverse cultural contexts—whether guided by abstract values or concrete situational cues. Method: This paper introduces NormAd, the first systematic evaluation framework for quantifying LLMs’ cultural adaptability across multi-granular cultural layers—from universal values to country-specific norms. It establishes NormAd-Eti, a cross-cultural etiquette benchmark comprising 2,600 situational prompts spanning 75 countries, and employs scenario-based reasoning evaluation, multi-level prompting strategies, and human baseline comparisons. Results: Experiments reveal that even the strongest models achieve <82% accuracy under explicit normative guidance (vs. >95% for humans), plummeting to <60% when provided only with abstract values and country identifiers (vs. >90% for humans). Notably, models exhibit pronounced bias against Global South cultures, underscoring critical gaps in culturally grounded reasoning.
This work addresses the challenge that existing cultural alignment methods struggle to simultaneously preserve the generalized cultural values of large language models and optimize downstream task performance, often suffering from cross-cultural interference. To this end, the authors propose CultureManager, a novel framework that introduces, for the first time, a task-aware cultural data synthesis mechanism coupled with a routing-based multi-cultural adapter architecture. The approach constructs culturally relevant data via web search, synthesizes it with task-format alignment, and employs dedicated cultural adapters with dynamic routing for modular management. Extensive experiments across ten national cultures and multiple culturally sensitive tasks demonstrate that CultureManager significantly outperforms prompt engineering and fine-tuning baselines, effectively mitigating cross-cultural conflicts while enhancing task adaptability.
This study addresses how human value is reconfigured within the language and translation industry amid the surge of automation and how this transformation influences translation pedagogy. Drawing on qualitative interviews with 29 industry stakeholders and integrating Chesterman’s framework of translation ethics with data from the LT-LiDER project, the research proposes “adaptability” as a central mediating concept linking human professional value with technological efficiency. Findings indicate that while technological efficiency has become a baseline expectation, human value is reaffirmed through professional judgment, oversight, and accountability embedded within automated workflows. The study demonstrates that automation does not replace but rather reshapes the translation value system, offering a new direction for translation education oriented toward human–machine collaboration.
This study addresses persistent challenges in cross-lingual translation of mathematical word problems by large language models, particularly insufficient cultural consistency, diversity compression, and misjudgment of regional context. Through the first large-scale corpus analysis, the authors systematically audit how Claude Opus 4, GPT-4.1, and Gemini 2.5 Pro handle culturally embedded entities—such as names, foods, and locations—when translating 60 English problems into seven languages spanning high- and low-resource settings. Combining human annotation with quantitative evaluation, they perform fine-grained coding of 6,489 cultural transformation instances, revealing widespread entropy collapse in diversity, surface-level token preferences, systematic regional misattribution, and frequent cross-cultural contamination errors (e.g., “Easter eggs used in Eid celebrations”). The work introduces the first fine-grained annotation framework for cultural translation, finding model agreement on transformation type in only 62.5% of cases and exact substitution alignment in just 33.5%.
This study systematically evaluates the generalization of commonsense knowledge in large language models across multilingual and multicultural contexts, with a particular focus on low-resource languages and underrepresented cultures. Building upon a human-curated extension of the BLEnD benchmark encompassing over 30 language–culture pairs, the evaluation features two tracks—short-answer and multiple-choice—and strictly enforces a zero-shot setting, prohibiting any training or fine-tuning on the benchmark data while allowing participation from any NLP system. As the first large-scale, purely evaluative benchmark for cross-cultural commonsense reasoning, the initiative attracted registrations from over 140 teams, with 62 submitting results. Analysis reveals that state-of-the-art approaches perform substantially worse on low-resource languages, highlighting critical challenges in cultural alignment and cross-cultural commonsense transfer.