The influence of missing data mechanisms and simple missing data handling techniques on fairness

📅 2025-03-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study systematically investigates how missing data mechanisms—specifically Missing Completely at Random (MCAR) and Missing at Random (MAR)—interact with common imputation methods (listwise deletion, mean/mode imputation, k-Nearest Neighbors imputation) to affect algorithmic fairness in machine learning. Using simulation experiments on three benchmark fairness datasets, we find that under MAR, simple strategies such as listwise deletion and mode imputation significantly improve fairness—reducing statistical bias by up to 42%—despite modest reductions in predictive accuracy; conversely, high-fidelity kNN imputation exacerbates group-level bias. To our knowledge, this is the first work to uncover an intrinsic link between missingness mechanisms and algorithmic fairness, challenging the prevailing assumption that complex imputation inherently yields fairer outcomes. Our findings introduce a mechanism-aware perspective for handling missing data in fair ML, offering both theoretical insight and actionable guidance for practitioners.

Technology Category

Machine Learning: Ethics, Bias, and FairnessComputer Vision: Bias, Fairness & PrivacyPhilosophy and Ethics of AI: Bias, Fairness & Equity

Application Category

User Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSocial Networks and Social Media: Fairness and bias in social network and social media analysisGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphs
📝 Abstract
Fairness of machine learning algorithms is receiving increasing attention, as such algorithms permeate the day-to-day aspects of our lives. One way in which bias can manifest in a dataset is through missing values. If data are missing, these data are often assumed to be missing completely randomly; in reality the propensity of data being missing is often tied to the demographic characteristics of individuals. There is limited research into how missing values and the handling thereof can impact the fairness of an algorithm. Most researchers either apply listwise deletion or tend to use the simpler methods of imputation (e.g. mean or mode) compared to the more advanced ones (e.g. multiple imputation); we therefore study the impact of the simpler methods on the fairness of algorithms. The starting point of the study is the mechanism of missingness, leading into how the missing data are processed and finally how this impacts fairness. Three popular datasets in the field of fairness are amputed in a simulation study. The results show that under certain scenarios the impact on fairness can be pronounced when the missingness mechanism is missing at random. Furthermore, elementary missing data handling techniques like listwise deletion and mode imputation can lead to higher fairness compared to more complex imputation methods like k-nearest neighbour imputation, albeit often at the cost of lower accuracy.
Problem

Research questions and friction points this paper is trying to address.

Impact of missing data mechanisms on algorithm fairness
Effect of simple missing data handling techniques on fairness
Comparison of fairness outcomes using different imputation methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

Study impact of simple missing data techniques
Analyze fairness effects of missingness mechanisms
Compare elementary vs complex imputation methods
A
Aeysha Bhatti
Department of Statistics and Actuarial Science, Stellenbosch University, Stellenbosch, 7602, South Africa
T
Trudie Sandrock
Department of Statistics and Actuarial Science, Stellenbosch University, Stellenbosch, 7602, South Africa
J
Johané Nienkemper-Swanepoel
Department of Statistics and Actuarial Science, Stellenbosch University, Stellenbosch, 7602, South Africa