🤖 AI Summary
This study investigates whether large language models (LLMs) retain nationality bias when explicit nationality labels are removed, yet culturally associated names remain. We propose a novel name-substitution–based bias evaluation framework, extending the Bias Benchmark for QA (BBQ) to better reflect real-world ambiguous contexts. Our method systematically measures bias persistence and reasoning accuracy across leading models—including those from OpenAI, Google, and Anthropic. Methodologically, we decouple cultural symbolism from explicit demographic markers for the first time, exposing how models respond to implicit cues. Results show that smaller-parameter models exhibit stronger bias and lower accuracy (e.g., Claude Haiku: 9% bias rate vs. Sonnet’s 3.5%). GPT-4o-mini retains errors in 76% of ambiguous cases—higher than GPT-4o’s 68%—indicating a positive correlation between model scale and bias robustness.
📝 Abstract
Large Language Models (LLMs) can exhibit latent biases towards specific nationalities even when explicit demographic markers are not present. In this work, we introduce a novel name-based benchmarking approach derived from the Bias Benchmark for QA (BBQ) dataset to investigate the impact of substituting explicit nationality labels with culturally indicative names, a scenario more reflective of real-world LLM applications. Our novel approach examines how this substitution affects both bias magnitude and accuracy across a spectrum of LLMs from industry leaders such as OpenAI, Google, and Anthropic. Our experiments show that small models are less accurate and exhibit more bias compared to their larger counterparts. For instance, on our name-based dataset and in the ambiguous context (where the correct choice is not revealed), Claude Haiku exhibited the worst stereotypical bias scores of 9%, compared to only 3.5% for its larger counterpart, Claude Sonnet, where the latter also outperformed it by 117.7% in accuracy. Additionally, we find that small models retain a larger portion of existing errors in these ambiguous contexts. For example, after substituting names for explicit nationality references, GPT-4o retains 68% of the error rate versus 76% for GPT-4o-mini, with similar findings for other model providers, in the ambiguous context. Our research highlights the stubborn resilience of biases in LLMs, underscoring their profound implications for the development and deployment of AI systems in diverse, global contexts.