๐ค AI Summary
This study addresses the critical gap in evaluating social bias in multilingual large language models, which has predominantly relied on English and thus fails to capture risks in cross-lingual deployment. By submitting 4,900 symmetric EnglishโSwahili prompt pairs to GPT-5.2 and Gemini 2.5 Flash, the authors systematically analyze bias across four dimensions: stereotyping, sentiment polarity, refusal behavior, and semantic consistency. Their findings reveal that bias does not merely transfer across languages but undergoes structural transformation. Notably, model refusals are heavily contingent on English surface forms, undermining the validity of monolingual audits. Empirical results show a 12-percentage-point disparity in stereotyping rates, a doubling of neutral sentiment responses for Gemini in Swahili, zero refusals by GPT-5.2 in Swahili compared to 169 in English, and semantic inconsistencies in over 55% of cross-lingual outputs.
๐ Abstract
Large language models are increasingly deployed in multilingual contexts, yet safety alignment and bias evaluation remain overwhelmingly English-centric. We investigate whether social biases generalise across languages by submitting 4,900 symmetric English--Swahili prompt pairs to GPT-5.2 and Gemini 2.5 Flash across nine demographic bias axes, yielding 19,600 completions evaluated for stereotype prevalence, sentiment, refusal behaviour, and cross-lingual semantic similarity. Our findings show that bias transforms rather than transfers: stereotype rates shifted by up to 12 percentage points on specific axes, Gemini's neutral-sentiment rate doubled in Swahili, and GPT-5.2 refused 169 prompts in English and zero in Swahili, consistent with refusal behaviour anchored to English-language surface forms at the behavioural level. Over 55% of prompt pairs produced semantically dissimilar completions across both models. These reinforce the idea that English-only bias audits do not produce adequate coverage for multilingual deployment.