Study of the influence of a biased database on the prediction of standard algorithms for selecting the best candidate for an interview

📅 2025-05-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates fairness degradation in AI-driven recruitment under dual bias—external discrimination and internal self-censorship—and evaluates the mitigating effect of resume anonymization. We propose the first systematic framework for synthesizing realistic, reproducible datasets that jointly encode both bias types. Using these data, we quantitatively assess fairness decay across five standard classifiers—logistic regression, decision trees, random forests, SVM, and XGBoost—under objective evaluation criteria. Results show an average 37% drop in top-candidate recall across all models on biased data; anonymization yields only marginal fairness improvement and fails to eliminate systemic bias rooted in historical inequities. Our core contributions are: (1) a transparent, reproducible dual-bias data generation framework; and (2) empirical evidence demonstrating the fundamental limitations of anonymization in algorithmic hiring, establishing a new benchmark for bias溯源 (origin tracing) and mitigation.

Technology Category

Computer Vision: Bias, Fairness & PrivacyMachine Learning: Ethics, Bias, and FairnessNatural Language Processing: Ethics — Bias, Fairness, Transparency & Privacy

Application Category

User Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSocial Networks and Social Media: Fairness and bias in social network and social media analysisSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Artificial intelligence is used at various stages of the recruitment process to automatically select the best candidate for a position, with companies guaranteeing unbiased recruitment. However, the algorithms used are either trained by humans or are based on learning from past experiences that were biased. In this article, we propose to generate data mimicking external (discrimination) and internal biases (self-censorship) in order to train five classic algorithms and to study the extent to which they do or do not find the best candidates according to objective criteria. In addition, we study the influence of the anonymisation of files on the quality of predictions.
Problem

Research questions and friction points this paper is trying to address.

Investigates biased database impact on interview candidate prediction algorithms
Examines AI recruitment bias from human-trained or historical data
Analyzes anonymization effect on prediction quality in candidate selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generates data mimicking external and internal biases
Trains five classic algorithms with biased data
Studies anonymisation impact on prediction quality
🔎 Similar Papers
2023-09-25ACM Transactions on Intelligent Systems and TechnologyCitations: 21
S
Shuyu Wang
Univ. Grenoble Alpes, CNRS, Grenoble INP∗, LJK, 38000 Grenoble, France
A
Ang'elique Saillet
Univ. Grenoble Alpes, CNRS, Grenoble INP∗, LJK, 38000 Grenoble, France
P
Philomene Le Gall
Univ. Grenoble Alpes, CNRS, Grenoble INP∗, LJK, 38000 Grenoble, France
A
Alain Lacroux
Universit´e Paris 1 Panth´eon-Sorbonne, PRISM, Paris, France
C
Christelle Martin-Lacroux
Univ. Grenoble Alpes, Grenoble INP∗, CERAG, 38000 Grenoble, France
V
Vincent Brault
Univ. Grenoble Alpes, CNRS, Grenoble INP∗, LJK, 38000 Grenoble, France