Score
Design, implement, and evaluate methods and workflows to detect, model, and replace missing entries in datasets, including single and multiple imputation, model-based or matrix-completion approaches, and strategies for handling missingness during preprocessing. This competence also covers assessing and selecting assumptions about the missingness mechanism (MCAR/MAR/MNAR), quantifying imputation uncertainty, and measuring the impact of imputation on downstream analyses and model performance.
To address missing data arising simultaneously from MCAR, MAR, and MNAR mechanisms in real-world scenarios, this paper proposes the first mechanism-adaptive, multimodal robust framework. Methodologically, it introduces the first systematic unification of all three missingness mechanisms, integrating causal inference, variational autoencoders, uncertainty modeling, and adversarial training to jointly enable missing pattern identification, dynamic mechanism discrimination, and end-to-end optimization. The framework supports heterogeneous real-world data—including tabular, time-series, and image modalities—thereby overcoming the restrictive MCAR-dominant assumption prevalent in prior work. Evaluated across 12 cross-domain benchmarks, it achieves an average 19.3% improvement in imputation accuracy and attains downstream classification and prediction performance comparable to that of models trained on complete data. The framework has been deployed in an industrial-grade data governance platform.
Conventional imputation methods (e.g., MICE, Amelia) assume missingness at random (MAR) and smooth data distributions, leading to model misspecification and bias under missing not at random (MNAR) mechanisms—particularly when missingness patterns are sparse. Method: We propose a novel graphical-model-based framework for characterizing the full-data law, introducing for the first time a recursive equation system that explicitly models both MNAR mechanisms and sparse missingness structures—without requiring MAR or parametric assumptions. Leveraging graphical model theory and Gibbs sampling, we design a chained multivariate imputation procedure conceptually analogous to MICE but with formal theoretical interpretability. Results: Empirical evaluation shows our method matches MICE’s performance under MAR settings, while significantly reducing bias and improving imputation accuracy under MNAR and high-dimensional sparse missingness scenarios.
This work addresses the lack of a unified theoretical framework for handling missing data, particularly under missing-not-at-random (MNAR) mechanisms where existing methods often fail to ensure consistent prediction. The authors propose a novel framework that explicitly distinguishes between two prediction objectives—depending on whether the observation indicators of variables are utilized—and introduces a fine-grained classification of missingness mechanisms accordingly. Building on this distinction, they establish conditions weaker than missing-at-random (MAR) under which consistent prediction remains achievable. By integrating probabilistic modeling, pattern-wise submodeling, and unconditional imputation, the framework supports a comprehensive prediction theory spanning model development, validation, and deployment. Empirical evaluations on both synthetic data and a real-world emergency trauma prediction task demonstrate that the proposed approach consistently achieves optimal predictive performance across diverse missingness mechanisms, thereby overcoming the limitations of conventional methods reliant on the MAR assumption.
The prevailing assumption that higher imputation accuracy necessarily improves downstream predictive performance lacks empirical validation. Method: We systematically evaluate the impact of imputation accuracy on prediction across 19 real and synthetic datasets, testing 12 imputation methods (e.g., mean, KNN, MICE, GAIN) combined with diverse linear and nonlinear predictors (e.g., XGBoost, MLP). Contribution/Results: Under expressive predictive models, gains in imputation accuracy yield negligible improvements in final prediction performance. Missingness indicators consistently and significantly enhance generalization across MCAR and multiple missingness mechanisms. Imputation accuracy only meaningfully affects prediction in linearly generated data—not in real-world datasets. These findings challenge the conventional “imputation-first” paradigm and advocate a prediction-oriented approach to missing-data handling. The study provides empirically grounded, efficient, and robust guidance for practical modeling, emphasizing task-relevant signal preservation over fidelity of imputed values.
Parameter estimation under missing data is doubly sensitive to both misspecification of the underlying data model and deviations from standard missingness mechanisms (MCAR/MAR/MNAR) or Huber-type contamination. Method: This paper proposes a robust M-estimation framework grounded in the Maximum Mean Discrepancy (MMD), leveraging kernel embeddings and functional analysis tools. It avoids explicit specification of either the missingness mechanism or the full-data distribution. Contribution/Results: The method achieves joint robustness against both types of misspecification—theoretically guaranteeing strong consistency and asymptotic normality under MCAR. It yields a decomposable, explicit error bound that cleanly separates model misspecification error from missingness-induced bias. Moreover, it maintains controlled estimation error under MNAR and Huber contamination. By circumventing stringent modeling assumptions, the approach significantly enhances robustness and reliability in practical applications.
Uncertainty quantification in missing data imputation is often overlooked, and the relationship between calibration quality and imputation accuracy remains poorly understood. This paper presents the first systematic empirical evaluation of six state-of-the-art imputation methods—statistical (MICE, SoftImpute), distribution-alignment (OT-Impute), and deep generative (GAIN, MIWAE, TabCSDI)—across multiple real-world datasets, under MCAR, MAR, and MNAR missingness mechanisms, and across varying missing rates. We propose a multi-path evaluation framework integrating repeated sampling variability, conditional distribution modeling, and predictive confidence quantification to rigorously assess uncertainty calibration. Results reveal that high imputation accuracy does not imply well-calibrated uncertainty estimates; significant trade-offs exist among accuracy, calibration fidelity, and computational efficiency across method categories. We identify several robust, reproducible configurations, providing actionable, evidence-based guidance for model selection in downstream machine learning and data cleaning tasks.
Current approaches to sample size calculation for clinical prediction models typically neglect the impact of missing data, often resulting in overfitting and poor calibration. This study is the first to integrate missing data mechanisms and handling strategies—such as multiple imputation—into a posterior-distribution-based sample size framework. Through simulation studies and Expected Value of Perfect Information (EVPI) analyses, the research quantifies how missingness affects model performance. Findings reveal that under common missing data scenarios, even when existing minimum sample size criteria are met, calibration slopes frequently fall below 0.9. In certain settings, nearly twice the conventional sample size is required to achieve performance comparable to that with complete data, underscoring both the necessity and feasibility of dynamically adjusting sample size requirements in the presence of missing data.
This study addresses the challenge of diminished stability and generalizability of clinical prediction models under complex missing data, where the impact of different imputation strategies remains unclear. Leveraging a real-world cardiac disease cohort, we simulated 18 distinct missingness mechanisms to systematically evaluate how multiple imputation, missForest, k-nearest neighbors (kNN) imputation, and complete-case analysis affect logistic regression model performance. Model assessment encompassed internal and external validation metrics including AUC, calibration slope, prediction error, and computational efficiency. Our work provides the first comprehensive comparison of imputation methods across diverse missing data patterns, revealing that kNN imputation demonstrates superior robustness—particularly under high missingness rates and complex missingness structures—while achieving excellent external generalizability and the lowest computational cost, making it especially suitable for large-scale clinical modeling.
Sequential multiple assignment randomized trials (SMARTs) present unique challenges for causal inference due to complex missing data patterns arising from response-dependent re-randomization. This study systematically evaluates the extent to which existing statistical methods address SMART-specific missingness through a narrative review and predefined secondary data extraction, complemented by an analysis of reporting practices in 30 empirical studies. It provides the first comprehensive synthesis of missing data methodologies tailored to SMART designs, revealing that only one of seven methodological papers fully accounts for all relevant missing data types. Among empirical studies, the median attrition rate was 18.1%, and merely 14% pre-specified sensitivity analyses for missing data, highlighting a substantial gap between methodological advances and their implementation in practice.
This study addresses a critical limitation of deterministic imputation methods based on minimizing mean squared error (MSE), which, despite yielding accurate point estimates, systematically underestimate data variability and thereby introduce bias into downstream statistics such as variance, correlation, and regression coefficients. To rectify this, the authors propose a stochastic imputation strategy that augments MSE-optimal predictions with random noise scaled to the residual error variance, thereby restoring the original distributional properties of the data. Through multivariate normal simulations, the work demonstrates for the first time the inadequacy of MSE as a sole imputation quality metric and reveals pervasive bias in widely used predictive imputation methods—including missForest, softImpute, and MICE. The proposed stochastic approach effectively eliminates this bias, ensuring statistical validity in subsequent analyses and advocating a paradigm shift from deterministic to stochastic imputation.