Retrieval Sensitivity to Identity Signals in Queries

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of dense retrievers to identity signals in queries, such as political ideology and dialect, which can induce unfair information filtering and exacerbate health disparities. To systematically quantify the independent effects of these signals on retrieval, this work constructs controlled synthetic datasets for comparison with real-world data. By employing lexical asymmetry controls, sparse baselines, and embedding linear probes, it effectively disentangles surface-level lexical confounds, demonstrating that the observed biases originate from deep semantic representations rather than mere word overlap. Experimental results reveal that all evaluated retrievers exhibit systematic bias toward accommodating users' political stances, while yielding significantly degraded performance on African American English compared to mainstream English. These findings highlight critical fairness risks inherent in current retrieval systems that urgently require mitigation.
📝 Abstract
Dense retrievers decide which documents reach users and the language models that use them, yet they are typically evaluated with neutral queries. We ask whether the identity signals that real users express in their queries---political ideology and dialect---bias what a retriever returns. We design evaluations in two domains, political news and consumer-health questions, each pairing a controlled synthetic set that varies only the identity signal with naturalistic queries. Across five dense retrievers and a sparse baseline, every retriever (i) retrieves articles that align with the query's own political lean and (ii) performs worse for questions written in African American Language (AAL) than in White Mainstream English (WME). Two analyses tie these gaps to queries' identity signals beyond surface vocabulary: partialling out an aggregate lexical-asymmetry score leaves the synthetic gaps largely intact, and linear probes recover lean and dialect from the retrievers' query embeddings beyond token-level features. Left unaddressed, such retrieval biases risk contributing to polarization and reinforcing the health disparities already faced by AAL speakers. Code is available at https://github.com/Andrewtcr/bias-ret.
Problem

Research questions and friction points this paper is trying to address.

dense retrievers
retrieval bias
identity signals
political ideology
dialect
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dense Retrievers
Identity Signals
Retrieval Bias
African American Language
Linear Probes
🔎 Similar Papers
No similar papers found.