missing data handling

Designs, builds, and evaluates methods for handling datasets with missing entries, including algorithms that impute missing values (single and stochastic multiple imputation), preserve or impute multivariate structures such as correlations, represent and generate missingness indicators, and jointly model value–missingness processes and missingness mechanisms. Implements identification-aware and identification-guided imputation and training-mask selection, generates multiple completed datasets or composite feature vectors, simulates missingness patterns, and quantifies impacts on bias, variance, uncertainty propagation, and recovery of target extrapolation distributions.

missingdatahandling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.58
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study systematically investigates how missing data mechanisms—specifically Missing Completely at Random (MCAR) and Missing at Random (MAR)—interact with common imputation methods (listwise deletion, mean/mode imputation, k-Nearest Neighbors imputation) to affect algorithmic fairness in machine learning. Using simulation experiments on three benchmark fairness datasets, we find that under MAR, simple strategies such as listwise deletion and mode imputation significantly improve fairness—reducing statistical bias by up to 42%—despite modest reductions in predictive accuracy; conversely, high-fidelity kNN imputation exacerbates group-level bias. To our knowledge, this is the first work to uncover an intrinsic link between missingness mechanisms and algorithmic fairness, challenging the prevailing assumption that complex imputation inherently yields fairer outcomes. Our findings introduce a mechanism-aware perspective for handling missing data in fair ML, offering both theoretical insight and actionable guidance for practitioners.

Comparison of fairness outcomes using different imputation methodsEffect of simple missing data handling techniques on fairnessImpact of missing data mechanisms on algorithm fairness

Multiple data-driven missing imputation

Jul 03, 2025
SK
Sergii Kavun
🏛️ Interregional Academy of Personnel Management | Luxena Ltd.

This paper addresses the adaptive imputation of short-to-moderate-length (1–5+ points) missing segments in univariate time series—occurring at arbitrary positions (beginning, middle, or end)—where local temporal characteristics (e.g., trend, volatility) vary significantly. To this end, we propose KZImputer, a novel method that dynamically selects the optimal imputation strategy based on both missing-segment location and local signal structure: extrapolative linear fitting for leading gaps, a hybrid of local mean and linear interpolation for interior gaps, and backward trend estimation for trailing gaps. KZImputer demonstrates robust performance even under high missingness rates (>50%), substantially outperforming conventional methods (e.g., linear, spline, and last-observation-carried-forward imputation). Extensive experiments show that it achieves state-of-the-art accuracy in MAE and RMSE, while preserving signal morphology more faithfully—as quantified by dynamic time warping (DTW) distance and spectral similarity. Notably, gains are most pronounced in sparse-data regimes, enhancing reliability for downstream analytical tasks.

Handling gaps at start, middle, or end of seriesImproving accuracy for high missingness rates (50%+)Imputing missing points in univariate time series data

Missing Value Knockoffs

Feb 26, 2022
DK
Deniz Koyuncu
🏛️ Rensselaer Polytechnic Institute

Existing variable selection methods struggle to control the false discovery rate (FDR) under missing data, while model-X knockoffs—though theoretically guaranteed to control FDR—cannot directly accommodate missing values. This work establishes, for the first time, the theoretical FDR controllability of knockoffs in the presence of missing data. We propose three novel paradigms: (i) posterior sampling-based imputation and knockoff reuse, (ii) knockoff generation restricted to observed variables only, and (iii) joint latent-variable imputation and knockoff construction. Our approaches integrate Bayesian posterior sampling, univariate imputation, and latent-variable modeling, and we rigorously prove that they satisfy FDR ≤ α under standard assumptions. Extensive experiments demonstrate precise FDR control across diverse missingness mechanisms (MCAR, MAR, MNAR), variable correlation structures, and sample sizes, while achieving high statistical power and substantially reduced computational complexity compared to existing alternatives.

Extending model-x knockoffs framework to handle missing dataPreserving false selection guarantees with imputation methodsReducing computational complexity for latent variable models

Beyond Accuracy: An Empirical Study of Uncertainty Estimation in Imputation

Nov 26, 2025
ZT
Zarin Tahia Hossain
🏛️ Western University

Uncertainty quantification in missing data imputation is often overlooked, and the relationship between calibration quality and imputation accuracy remains poorly understood. This paper presents the first systematic empirical evaluation of six state-of-the-art imputation methods—statistical (MICE, SoftImpute), distribution-alignment (OT-Impute), and deep generative (GAIN, MIWAE, TabCSDI)—across multiple real-world datasets, under MCAR, MAR, and MNAR missingness mechanisms, and across varying missing rates. We propose a multi-path evaluation framework integrating repeated sampling variability, conditional distribution modeling, and predictive confidence quantification to rigorously assess uncertainty calibration. Results reveal that high imputation accuracy does not imply well-calibrated uncertainty estimates; significant trade-offs exist among accuracy, calibration fidelity, and computational efficiency across method categories. We identify several robust, reproducible configurations, providing actionable, evidence-based guidance for model selection in downstream machine learning and data cleaning tasks.

Analyzes trade-offs between accuracy, calibration, and runtime in imputationCompares statistical, distribution alignment, and deep generative imputation techniquesEvaluates uncertainty estimation reliability in imputation methods

Imputation for prediction: beware of diminishing returns

Jul 29, 2024
ML
Marine Le Morvan
🏛️ Inria

The prevailing assumption that higher imputation accuracy necessarily improves downstream predictive performance lacks empirical validation. Method: We systematically evaluate the impact of imputation accuracy on prediction across 19 real and synthetic datasets, testing 12 imputation methods (e.g., mean, KNN, MICE, GAIN) combined with diverse linear and nonlinear predictors (e.g., XGBoost, MLP). Contribution/Results: Under expressive predictive models, gains in imputation accuracy yield negligible improvements in final prediction performance. Missingness indicators consistently and significantly enhance generalization across MCAR and multiple missingness mechanisms. Imputation accuracy only meaningfully affects prediction in linearly generated data—not in real-world datasets. These findings challenge the conventional “imputation-first” paradigm and advocate a prediction-oriented approach to missing-data handling. The study provides empirically grounded, efficient, and robust guidance for practical modeling, emphasizing task-relevant signal preservation over fidelity of imputed values.

Assesses imputation's role in real-data predictive accuracy.Explores conditions favoring simple over complex imputation.Investigates impact of advanced imputation on predictions.

Latest Papers

What's happening recently
View more

This work addresses the lack of a unified theoretical framework for handling missing data, particularly under missing-not-at-random (MNAR) mechanisms where existing methods often fail to ensure consistent prediction. The authors propose a novel framework that explicitly distinguishes between two prediction objectives—depending on whether the observation indicators of variables are utilized—and introduces a fine-grained classification of missingness mechanisms accordingly. Building on this distinction, they establish conditions weaker than missing-at-random (MAR) under which consistent prediction remains achievable. By integrating probabilistic modeling, pattern-wise submodeling, and unconditional imputation, the framework supports a comprehensive prediction theory spanning model development, validation, and deployment. Empirical evaluations on both synthetic data and a real-world emergency trauma prediction task demonstrate that the proposed approach consistently achieves optimal predictive performance across diverse missingness mechanisms, thereby overcoming the limitations of conventional methods reliant on the MAR assumption.

missing datamissingness mechanismnon-MAR

This study addresses a critical limitation of deterministic imputation methods based on minimizing mean squared error (MSE), which, despite yielding accurate point estimates, systematically underestimate data variability and thereby introduce bias into downstream statistics such as variance, correlation, and regression coefficients. To rectify this, the authors propose a stochastic imputation strategy that augments MSE-optimal predictions with random noise scaled to the residual error variance, thereby restoring the original distributional properties of the data. Through multivariate normal simulations, the work demonstrates for the first time the inadequacy of MSE as a sole imputation quality metric and reveals pervasive bias in widely used predictive imputation methods—including missForest, softImpute, and MICE. The proposed stochastic approach effectively eliminates this bias, ensuring statistical validity in subsequent analyses and advocating a paradigm shift from deterministic to stochastic imputation.

data variabilitydownstream analysisimputation bias

This study addresses selection bias arising from missing data in graphical models by proposing an intervention-based framework for identification and estimation. Treating missingness indicators as intervenable variables, the authors introduce a novel tree-based identification algorithm to explicitly characterize the propagation pathways of selection bias and leverage do-calculus to assess the identifiability of target functionals. Building on this foundation, they develop a recursive inverse probability weighting approach to construct efficient estimating equations that jointly model the missingness mechanism and parameters of interest. The accompanying R package, flexMissing, enables end-to-end analysis, and both simulation studies and real-data applications demonstrate that the proposed method accurately determines identifiability and yields robust estimates.

graphical modelsidentifiabilityintervention

This study addresses the challenge of diminished stability and generalizability of clinical prediction models under complex missing data, where the impact of different imputation strategies remains unclear. Leveraging a real-world cardiac disease cohort, we simulated 18 distinct missingness mechanisms to systematically evaluate how multiple imputation, missForest, k-nearest neighbors (kNN) imputation, and complete-case analysis affect logistic regression model performance. Model assessment encompassed internal and external validation metrics including AUC, calibration slope, prediction error, and computational efficiency. Our work provides the first comprehensive comparison of imputation methods across diverse missing data patterns, revealing that kNN imputation demonstrates superior robustness—particularly under high missingness rates and complex missingness structures—while achieving excellent external generalizability and the lowest computational cost, making it especially suitable for large-scale clinical modeling.

clinical prediction modelsimputation methodsmissing data

Current approaches to sample size calculation for clinical prediction models typically neglect the impact of missing data, often resulting in overfitting and poor calibration. This study is the first to integrate missing data mechanisms and handling strategies—such as multiple imputation—into a posterior-distribution-based sample size framework. Through simulation studies and Expected Value of Perfect Information (EVPI) analyses, the research quantifies how missingness affects model performance. Findings reveal that under common missing data scenarios, even when existing minimum sample size criteria are met, calibration slopes frequently fall below 0.9. In certain settings, nearly twice the conventional sample size is required to achieve performance comparable to that with complete data, underscoring both the necessity and feasibility of dynamically adjusting sample size requirements in the presence of missing data.

clinical prediction modelsmissing datamodel calibration

Hot Scholars

YZ

Youran Zhou

PhD Student, Deakin University
missing dataimputationmissing mechanism
SA

Sunil Aryal

Deakin University Australia
Data miningMachine learning
TC

Tianlong Chen

Assistant Professor, CS@UNC Chapel Hill; Chief AI Scientist, hireEZ
Machine LearningAI4ScienceComputer VisionSparsity
AP

Antonio Punzo

Full Professor of Statistics, University of Catania
Mixture ModelsHidden Markov ModelsHeavy-tailed DistributionsSerial Dependence