Score
Design and implement metrics and evaluation protocols that measure how well missing data are detected and imputed, including accuracy of reconstructed values, calibration and uncertainty, and robustness across realistic missingness mechanisms. Build comparative analyses and downstream-task assessments that quantify the impact of imputation methods and validate whether imputation preserves or correctly identifies meaningful patterns of missingness.
Structural missingness patterns in high-dimensional data frequently introduce analytical bias, yet existing visualization methods lack both interpretability and scalability. This paper introduces the first explainable quality metric framework specifically designed for structural missingness, modeling missing patterns as multidimensional quantitative indicators—including pattern sparsity, dimensional coupling strength, and temporal regularity. We further propose a metric-driven visualization encoding and interactive analysis framework, enabling efficient exploration of large-scale, high-dimensional datasets. Experiments on real-world gait monitoring data demonstrate that our approach significantly improves structural missingness identification efficiency (3.2× faster than baseline methods) and diagnostic depth (revealing seven novel latent missing patterns). The framework delivers interpretable, actionable visual analytics support for data quality assessment and governance decision-making.
Uncertainty quantification in missing data imputation is often overlooked, and the relationship between calibration quality and imputation accuracy remains poorly understood. This paper presents the first systematic empirical evaluation of six state-of-the-art imputation methods—statistical (MICE, SoftImpute), distribution-alignment (OT-Impute), and deep generative (GAIN, MIWAE, TabCSDI)—across multiple real-world datasets, under MCAR, MAR, and MNAR missingness mechanisms, and across varying missing rates. We propose a multi-path evaluation framework integrating repeated sampling variability, conditional distribution modeling, and predictive confidence quantification to rigorously assess uncertainty calibration. Results reveal that high imputation accuracy does not imply well-calibrated uncertainty estimates; significant trade-offs exist among accuracy, calibration fidelity, and computational efficiency across method categories. We identify several robust, reproducible configurations, providing actionable, evidence-based guidance for model selection in downstream machine learning and data cleaning tasks.
Missing data are prevalent in scientific research and public health, often leading to analytical bias and compounded by a lack of effective tools for diagnosing missingness mechanisms and comparing imputation methods. This work proposes a method-agnostic visual analytics dashboard that integrates mainstream imputation techniques—including MICE, random forests, XGBoost, and kNN—and employs coordinated views such as heatmaps, co-missingness summaries, distribution diagnostics, and error metrics (e.g., MAE, RMSE) to facilitate missing pattern recognition, cross-method comparison, and downstream impact analysis. Innovatively, it introduces a geographically and socioeconomically informed gKNN algorithm to enable source-based visual accountability. Case studies demonstrate that the system effectively supports users in selecting imputation strategies, identifying sensitive variables, assessing model robustness, and substantially reducing the cognitive burden associated with switching between methods.
Evaluating imputation methods without ground-truth complete data remains challenging, as conventional metrics (e.g., RMSE) are misleading under realistic missingness mechanisms. Method: This paper proposes an unsupervised scoring framework based on the energy score, which constructs validation sets via controlled artificial masking and evaluates imputations by their ability to reproduce the underlying data distribution—under the Missing at Random (MAR) assumption. Contribution/Results: The framework is the first to explicitly model missingness mechanisms tailored to data distribution characteristics, ensuring scoring consistency with downstream task performance. Experiments on both synthetic and real-world datasets demonstrate its robustness in discriminating among imputation algorithms. Theoretically grounded and empirically validated, the approach offers both statistical soundness and practical deployability.
To address missing data arising simultaneously from MCAR, MAR, and MNAR mechanisms in real-world scenarios, this paper proposes the first mechanism-adaptive, multimodal robust framework. Methodologically, it introduces the first systematic unification of all three missingness mechanisms, integrating causal inference, variational autoencoders, uncertainty modeling, and adversarial training to jointly enable missing pattern identification, dynamic mechanism discrimination, and end-to-end optimization. The framework supports heterogeneous real-world data—including tabular, time-series, and image modalities—thereby overcoming the restrictive MCAR-dominant assumption prevalent in prior work. Evaluated across 12 cross-domain benchmarks, it achieves an average 19.3% improvement in imputation accuracy and attains downstream classification and prediction performance comparable to that of models trained on complete data. The framework has been deployed in an industrial-grade data governance platform.
Real-world missing data imputation urgently requires a comprehensive evaluation framework beyond point-estimation metrics such as RMSE. This paper reframes imputation as a distributional forecasting task and introduces *Imputation Scores*—a novel metric quantifying how well imputed values preserve the underlying data distribution. For the first time, we systematically evaluate over a dozen imputation methods under realistic missingness mechanisms—not merely synthetic MCAR or MAR settings. Experiments span numerical and mixed-type datasets, employing the widely adopted iterative multiple imputation framework implemented in the *mice* R package, with unified benchmarking across both synthetic and real-world missing-data scenarios. Results demonstrate that iterative methods—particularly *mice*—significantly outperform single-imputation and deep learning approaches in distributional fidelity and robustness. Our framework provides practitioners with reproducible, interpretable criteria for method selection in practical applications. (149 words)
This work addresses the lack of a unified theoretical framework for handling missing data, particularly under missing-not-at-random (MNAR) mechanisms where existing methods often fail to ensure consistent prediction. The authors propose a novel framework that explicitly distinguishes between two prediction objectives—depending on whether the observation indicators of variables are utilized—and introduces a fine-grained classification of missingness mechanisms accordingly. Building on this distinction, they establish conditions weaker than missing-at-random (MAR) under which consistent prediction remains achievable. By integrating probabilistic modeling, pattern-wise submodeling, and unconditional imputation, the framework supports a comprehensive prediction theory spanning model development, validation, and deployment. Empirical evaluations on both synthetic data and a real-world emergency trauma prediction task demonstrate that the proposed approach consistently achieves optimal predictive performance across diverse missingness mechanisms, thereby overcoming the limitations of conventional methods reliant on the MAR assumption.
Existing approaches to surrogate marker evaluation often ignore missing data and rely on complete-case analyses, which can introduce bias and reduce statistical efficiency. This work proposes a unified framework that, for the first time, systematically incorporates missing data correction into surrogate marker validation by integrating inverse probability weighting (IPW) with semiparametric maximum likelihood estimation (SMLE). The framework ensures robustness, computational tractability, and high statistical efficiency under both nonparametric and parametric settings. An accompanying R package, MissSurrogate, implements the proposed methodology. Simulation studies demonstrate that the method remains unbiased and achieves near-full-sample efficiency across various missing data mechanisms. Its practical utility is further illustrated through an application to a diabetes clinical trial.
Current approaches to sample size calculation for clinical prediction models typically neglect the impact of missing data, often resulting in overfitting and poor calibration. This study is the first to integrate missing data mechanisms and handling strategies—such as multiple imputation—into a posterior-distribution-based sample size framework. Through simulation studies and Expected Value of Perfect Information (EVPI) analyses, the research quantifies how missingness affects model performance. Findings reveal that under common missing data scenarios, even when existing minimum sample size criteria are met, calibration slopes frequently fall below 0.9. In certain settings, nearly twice the conventional sample size is required to achieve performance comparable to that with complete data, underscoring both the necessity and feasibility of dynamically adjusting sample size requirements in the presence of missing data.
This study addresses the issue of metric failure and result bias caused by missing data in performance portability evaluations for high-performance computing (HPC). We systematically assess the effectiveness of multiple imputation algorithms in handling missing performance efficiency data. Leveraging real-world HPC benchmark datasets, we conduct an in-depth comparative analysis of these methods across three dimensions: accuracy, robustness, and computational cost, elucidating their applicable scenarios and inherent trade-offs. This work delineates the boundaries of strengths and limitations among different imputation strategies, providing concrete practical guidance and methodological support for mitigating the impact of missing data on performance portability analysis.