🤖 AI Summary
This study investigates fairness degradation in AI-driven recruitment under dual bias—external discrimination and internal self-censorship—and evaluates the mitigating effect of resume anonymization. We propose the first systematic framework for synthesizing realistic, reproducible datasets that jointly encode both bias types. Using these data, we quantitatively assess fairness decay across five standard classifiers—logistic regression, decision trees, random forests, SVM, and XGBoost—under objective evaluation criteria. Results show an average 37% drop in top-candidate recall across all models on biased data; anonymization yields only marginal fairness improvement and fails to eliminate systemic bias rooted in historical inequities. Our core contributions are: (1) a transparent, reproducible dual-bias data generation framework; and (2) empirical evidence demonstrating the fundamental limitations of anonymization in algorithmic hiring, establishing a new benchmark for bias溯源 (origin tracing) and mitigation.
📝 Abstract
Artificial intelligence is used at various stages of the recruitment process to automatically select the best candidate for a position, with companies guaranteeing unbiased recruitment. However, the algorithms used are either trained by humans or are based on learning from past experiences that were biased. In this article, we propose to generate data mimicking external (discrimination) and internal biases (self-censorship) in order to train five classic algorithms and to study the extent to which they do or do not find the best candidates according to objective criteria. In addition, we study the influence of the anonymisation of files on the quality of predictions.