Score
Designs, builds, and evaluates methods for handling datasets with missing entries, including algorithms that impute missing values (single and stochastic multiple imputation), preserve or impute multivariate structures such as correlations, represent and generate missingness indicators, and jointly model value–missingness processes and missingness mechanisms. Implements identification-aware and identification-guided imputation and training-mask selection, generates multiple completed datasets or composite feature vectors, simulates missingness patterns, and quantifies impacts on bias, variance, uncertainty propagation, and recovery of target extrapolation distributions.
This study systematically investigates how missing data mechanisms—specifically Missing Completely at Random (MCAR) and Missing at Random (MAR)—interact with common imputation methods (listwise deletion, mean/mode imputation, k-Nearest Neighbors imputation) to affect algorithmic fairness in machine learning. Using simulation experiments on three benchmark fairness datasets, we find that under MAR, simple strategies such as listwise deletion and mode imputation significantly improve fairness—reducing statistical bias by up to 42%—despite modest reductions in predictive accuracy; conversely, high-fidelity kNN imputation exacerbates group-level bias. To our knowledge, this is the first work to uncover an intrinsic link between missingness mechanisms and algorithmic fairness, challenging the prevailing assumption that complex imputation inherently yields fairer outcomes. Our findings introduce a mechanism-aware perspective for handling missing data in fair ML, offering both theoretical insight and actionable guidance for practitioners.
This paper addresses the adaptive imputation of short-to-moderate-length (1–5+ points) missing segments in univariate time series—occurring at arbitrary positions (beginning, middle, or end)—where local temporal characteristics (e.g., trend, volatility) vary significantly. To this end, we propose KZImputer, a novel method that dynamically selects the optimal imputation strategy based on both missing-segment location and local signal structure: extrapolative linear fitting for leading gaps, a hybrid of local mean and linear interpolation for interior gaps, and backward trend estimation for trailing gaps. KZImputer demonstrates robust performance even under high missingness rates (>50%), substantially outperforming conventional methods (e.g., linear, spline, and last-observation-carried-forward imputation). Extensive experiments show that it achieves state-of-the-art accuracy in MAE and RMSE, while preserving signal morphology more faithfully—as quantified by dynamic time warping (DTW) distance and spectral similarity. Notably, gains are most pronounced in sparse-data regimes, enhancing reliability for downstream analytical tasks.
Existing variable selection methods struggle to control the false discovery rate (FDR) under missing data, while model-X knockoffs—though theoretically guaranteed to control FDR—cannot directly accommodate missing values. This work establishes, for the first time, the theoretical FDR controllability of knockoffs in the presence of missing data. We propose three novel paradigms: (i) posterior sampling-based imputation and knockoff reuse, (ii) knockoff generation restricted to observed variables only, and (iii) joint latent-variable imputation and knockoff construction. Our approaches integrate Bayesian posterior sampling, univariate imputation, and latent-variable modeling, and we rigorously prove that they satisfy FDR ≤ α under standard assumptions. Extensive experiments demonstrate precise FDR control across diverse missingness mechanisms (MCAR, MAR, MNAR), variable correlation structures, and sample sizes, while achieving high statistical power and substantially reduced computational complexity compared to existing alternatives.
Uncertainty quantification in missing data imputation is often overlooked, and the relationship between calibration quality and imputation accuracy remains poorly understood. This paper presents the first systematic empirical evaluation of six state-of-the-art imputation methods—statistical (MICE, SoftImpute), distribution-alignment (OT-Impute), and deep generative (GAIN, MIWAE, TabCSDI)—across multiple real-world datasets, under MCAR, MAR, and MNAR missingness mechanisms, and across varying missing rates. We propose a multi-path evaluation framework integrating repeated sampling variability, conditional distribution modeling, and predictive confidence quantification to rigorously assess uncertainty calibration. Results reveal that high imputation accuracy does not imply well-calibrated uncertainty estimates; significant trade-offs exist among accuracy, calibration fidelity, and computational efficiency across method categories. We identify several robust, reproducible configurations, providing actionable, evidence-based guidance for model selection in downstream machine learning and data cleaning tasks.
The prevailing assumption that higher imputation accuracy necessarily improves downstream predictive performance lacks empirical validation. Method: We systematically evaluate the impact of imputation accuracy on prediction across 19 real and synthetic datasets, testing 12 imputation methods (e.g., mean, KNN, MICE, GAIN) combined with diverse linear and nonlinear predictors (e.g., XGBoost, MLP). Contribution/Results: Under expressive predictive models, gains in imputation accuracy yield negligible improvements in final prediction performance. Missingness indicators consistently and significantly enhance generalization across MCAR and multiple missingness mechanisms. Imputation accuracy only meaningfully affects prediction in linearly generated data—not in real-world datasets. These findings challenge the conventional “imputation-first” paradigm and advocate a prediction-oriented approach to missing-data handling. The study provides empirically grounded, efficient, and robust guidance for practical modeling, emphasizing task-relevant signal preservation over fidelity of imputed values.
This work addresses the lack of a unified theoretical framework for handling missing data, particularly under missing-not-at-random (MNAR) mechanisms where existing methods often fail to ensure consistent prediction. The authors propose a novel framework that explicitly distinguishes between two prediction objectives—depending on whether the observation indicators of variables are utilized—and introduces a fine-grained classification of missingness mechanisms accordingly. Building on this distinction, they establish conditions weaker than missing-at-random (MAR) under which consistent prediction remains achievable. By integrating probabilistic modeling, pattern-wise submodeling, and unconditional imputation, the framework supports a comprehensive prediction theory spanning model development, validation, and deployment. Empirical evaluations on both synthetic data and a real-world emergency trauma prediction task demonstrate that the proposed approach consistently achieves optimal predictive performance across diverse missingness mechanisms, thereby overcoming the limitations of conventional methods reliant on the MAR assumption.
This study addresses a critical limitation of deterministic imputation methods based on minimizing mean squared error (MSE), which, despite yielding accurate point estimates, systematically underestimate data variability and thereby introduce bias into downstream statistics such as variance, correlation, and regression coefficients. To rectify this, the authors propose a stochastic imputation strategy that augments MSE-optimal predictions with random noise scaled to the residual error variance, thereby restoring the original distributional properties of the data. Through multivariate normal simulations, the work demonstrates for the first time the inadequacy of MSE as a sole imputation quality metric and reveals pervasive bias in widely used predictive imputation methods—including missForest, softImpute, and MICE. The proposed stochastic approach effectively eliminates this bias, ensuring statistical validity in subsequent analyses and advocating a paradigm shift from deterministic to stochastic imputation.
This study addresses selection bias arising from missing data in graphical models by proposing an intervention-based framework for identification and estimation. Treating missingness indicators as intervenable variables, the authors introduce a novel tree-based identification algorithm to explicitly characterize the propagation pathways of selection bias and leverage do-calculus to assess the identifiability of target functionals. Building on this foundation, they develop a recursive inverse probability weighting approach to construct efficient estimating equations that jointly model the missingness mechanism and parameters of interest. The accompanying R package, flexMissing, enables end-to-end analysis, and both simulation studies and real-data applications demonstrate that the proposed method accurately determines identifiability and yields robust estimates.
This study addresses the challenge of diminished stability and generalizability of clinical prediction models under complex missing data, where the impact of different imputation strategies remains unclear. Leveraging a real-world cardiac disease cohort, we simulated 18 distinct missingness mechanisms to systematically evaluate how multiple imputation, missForest, k-nearest neighbors (kNN) imputation, and complete-case analysis affect logistic regression model performance. Model assessment encompassed internal and external validation metrics including AUC, calibration slope, prediction error, and computational efficiency. Our work provides the first comprehensive comparison of imputation methods across diverse missing data patterns, revealing that kNN imputation demonstrates superior robustness—particularly under high missingness rates and complex missingness structures—while achieving excellent external generalizability and the lowest computational cost, making it especially suitable for large-scale clinical modeling.
Current approaches to sample size calculation for clinical prediction models typically neglect the impact of missing data, often resulting in overfitting and poor calibration. This study is the first to integrate missing data mechanisms and handling strategies—such as multiple imputation—into a posterior-distribution-based sample size framework. Through simulation studies and Expected Value of Perfect Information (EVPI) analyses, the research quantifies how missingness affects model performance. Findings reveal that under common missing data scenarios, even when existing minimum sample size criteria are met, calibration slopes frequently fall below 0.9. In certain settings, nearly twice the conventional sample size is required to achieve performance comparable to that with complete data, underscoring both the necessity and feasibility of dynamically adjusting sample size requirements in the presence of missing data.