🤖 AI Summary
This study investigates whether large language models (LLMs) can accurately represent cultural variations in global social norms. Using the GSEN benchmark, we evaluate four mainstream LLMs via prompt engineering on norm perceptions across 90 societies and 150 scenarios, comparing their outputs against large-scale cross-national survey data. This work provides the first quantification of systematic biases in LLMs' cross-cultural understanding: models significantly underestimate the magnitude of cultural differences and mischaracterize their distributional patterns, yielding relatively accurate estimates only for vulgarity-related scenarios and developed societies. Furthermore, multilingual prompting strategies offer limited mitigation. By revealing fundamental limitations in LLMs' cultural cognition, this research provides critical empirical evidence for developing equitably aligned, cross-culturally competent AI systems.
📝 Abstract
A key aspect of culture is a society's norms about everyday behavior. How accurately do large language models (LLMs) represent cultural differences in such norms? To answer this question we used the recent Global Study of Everyday Norms (GSEN), which collected ratings of 150 scenarios in 90 societies, as the human benchmark. We prompted GPT-5 to estimate each society's average rating for every scenario, and later repeated the benchmark in three other LLMs: GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro. Compared to GSEN estimates, all four LLMs misrepresented cultural variation in two ways. First, they greatly underestimated its magnitude, estimating differences between societies to be, on average, less than half their measured size. Second, for many scenarios the LLMs poorly identified the pattern of variation, that is, which societies judged the behavior less acceptable and which societies judged it more acceptable. The pattern of variation was identified better for scenarios that elicit concerns about vulgarity, especially scenarios involving kissing and flirting. We also found that norms in more developed societies tended to be estimated somewhat more accurately, and that prompting in local survey languages rather than English produced only a modest improvement in accuracy. Local-language prompting also reduced, but did not remove, the underestimation of between-society differences. Cultural differences in everyday norms are only weakly and unevenly represented by LLMs.