STEREODISCO: Discovering Stereotypicality in LLMs

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of existing research in capturing the full structure of stereotypes embedded in large language models (LLMs), which stems from reliance on a narrow set of social-psychological semantic dimensions. To overcome this, the authors propose STEREODISCO, a novel framework that systematically integrates semantic differential scaling, WordNet antonym pairs, neural probing, and statistical testing to construct and identify stereotypical associations across approximately 2,000 semantic axes within LLM activation spaces. Validated on Llama-3-8B-Instruct and Mistral-7B-Instruct, the approach uncovers previously underexplored stereotype dimensions—such as humble–proud and narrow-minded–open-minded—and reveals high inter-model consistency in stereotype scoring that nonetheless significantly diverges from human judgments. These newly identified associations are further corroborated by human annotation.
📝 Abstract
LLMs encode, convey, and perpetuate stereotypes. Prior computational research focuses on a small set of semantic axes investigated in social psychology, and operates on word embeddings produced by language models, leaving open which other semantic axes carry stereotypical associations in LLMs and how LLMs internally represent such axes. We introduce STEREODISCO, a framework that adapts the semantic differential method (Osgood et al., 1957) to the systematic study of stereotypes in LLM internal representations. STEREODISCO constructs approx. 2,000 candidate semantic axes from WordNet antonym synsets, recovers each as a geometric axis in the LLM's activation space via probing, and identifies stereotypical axes via a statistical test over concept projections. As a case study, we apply STEREODISCO to social group stereotypes with LLAMA-3-8B-INSTRUCT and MISTRAL-7B-INSTRUCT. We find that the two LLMs agree with each other on social group ratings more than with humans, suggesting that LLM-encoded stereotype content diverges from that documented in social psychology. We also discover stereotypical axes not investigated in prior work -- including humble vs. proud, narrow-minded vs. broad-minded, and cowardly vs. brave, which human annotators independently confirm.
Problem

Research questions and friction points this paper is trying to address.

stereotypes
large language models
semantic axes
internal representations
bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

STEREODISCO
semantic axes
stereotype discovery
activation space probing
large language models
🔎 Similar Papers
No similar papers found.