🤖 AI Summary
This study addresses the underexplored issue of regional bias in large language models (LLMs) and its potential influence on social cognition and individual decision-making. The authors propose the Stereotypes-to-Decisions (S2D) framework, establishing the first comprehensive evaluation system covering all 34 provincial-level administrative regions of China. By linking abstract stereotype dimensions—such as warmth and competence—to concrete social decision tasks in education, employment, and social interaction, the study conducts a multidimensional assessment of six mainstream LLMs under both Chinese and English prompts. Results reveal consistent and significant regional biases in model outputs, closely aligned with regional economic and digital development indicators. These biases exhibit human-like yet asymmetric stereotyping patterns, underscoring the socio-structural origins of bias embedded in large language models.
📝 Abstract
Regional bias in large language models (LLMs) may shape both perceptions of regional groups and decisions about individuals from different regions. Yet existing studies often examine these manifestations separately, leaving their structure and consequences unclear. We introduce Stereotypes-to-Decisions (S2D), a systematic framework evaluating regional bias from abstract stereotypes to concrete social decisions. Covering all 34 provincial-level administrative regions of China, S2D evaluates six LLMs using stereotype ratings of Warmth (perceived friendliness and trustworthiness) and Competence (perceived capability and intelligence), along with paired-choice tasks across Education, Occupation, and Social Interaction. Results reveal substantial regional differences in regional scores, with considerable agreement across models, especially for Competence and Occupation decisions. Furthermore, these patterns are associated with regional economic and digital development indicators and display mixed human-like stereotypes, with some regions rated highly on one dimension but poorly on the other. They also remain largely stable across Chinese and English prompts. Overall, our findings show that regional bias in LLMs is prevalent, systematic, and consequential, motivating more regionally aware evaluation and mitigation.