๐ค AI Summary
This paper addresses two critical gaps in NLP fairness research: (1) the latent cultural biases exhibited by large language models (LLMs) in cross-cultural contexts, and (2) the widespread exclusion of real-world stakeholders from bias assessment. Through a meta-analysis of 20 recent (2025) NLP papers on cultural bias, we identify a foundational flawโoverreliance on static dataset-based evaluation detached from sociocultural context. To rectify this, we propose a novel methodology grounded in *stakeholder embedding*, introducing the first conceptual framework for cultural bias that prioritizes situated social impact over technical metrics. Building on this, we develop an actionable conceptualization guide and a structured social harm assessment pathway. This work establishes the first methodological benchmark for NLP fairness research explicitly designed for cross-cultural settings and integrating socio-technical perspectives.
๐ Abstract
Research has shown that while large language models (LLMs) can generate their responses based on cultural context, they are not perfect and tend to generalize across cultures. However, when evaluating the cultural bias of a language technology on any dataset, researchers may choose not to engage with stakeholders actually using that technology in real life, which evades the very fundamental problem they set out to address.
Inspired by the work done by arXiv:2005.14050v2, I set out to analyse recent literature about identifying and evaluating cultural bias in Natural Language Processing (NLP). I picked out 20 papers published in 2025 about cultural bias and came up with a set of observations to allow NLP researchers in the future to conceptualize bias concretely and evaluate its harms effectively. My aim is to advocate for a robust assessment of the societal impact of language technologies exhibiting cross-cultural bias.