Imputing With Predictive Mean Matching Can Be Severely Biased When Values Are Missing At Random

📅 2025-06-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study identifies a severe and systematic estimation bias in predictive mean matching (PMM) under the missing-at-random (MAR) mechanism. Through theoretical analysis and extensive simulation experiments, we demonstrate that when an observed covariate (X) strongly predicts the missingness probability of outcome (Y) and is highly correlated with (Y), PMM yields regression slope estimates biased by up to 80%. This bias persists even as the (X)–(Y) correlation weakens and only approaches zero under large samples ((n = 1{,}000)) and the more restrictive missing-completely-at-random (MCAR) assumption. Compared to alternative imputation methods, PMM exhibits greater sensitivity to the missingness mechanism and requires substantially larger sample sizes to achieve acceptable accuracy. Our work provides the first systematic characterization of how PMM bias varies with missingness patterns and sample size. These findings challenge the widespread recommendation of PMM as a default imputation method and deliver critical methodological warnings for applied missing-data analysis.

Technology Category

Reasoning under Uncertainty: Relational Probabilistic ModelsMachine Learning: Ensemble MethodsData Mining & Knowledge Management: Recommender Systems

Application Category

User Modeling, Personalization and Recommendation: Practical large-scale studies of user experienceGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methods
📝 Abstract
Predictive mean matching (PMM) is a popular imputation strategy that imputes missing values by borrowing observed values from other cases with similar expectations. We show that, unlike other imputation strategies, PMM is not guaranteed to be consistent -- and in fact can be severely biased -- when values are missing at random (when the probability a value is missing depends only on values that are observed). We demonstrate the bias in a simple situation where a complete variable $X$ is both strongly correlated with $Y$ and strongly predictive of whether $Y$ is missing. The bias in the estimated regression slope can be as large as 80 percent, and persists even when we reduce the correlation between $X$ and $Y$. To make the bias vanish, the sample must be large ($n$=1,000) emph{and} $Y$ values must be missing independently of $X$ (i.e., missing completely at random). Compared to other imputation methods, it seems that PMM requires larger samples and is more sensitive to the pattern of missing values. We cannot recommend PMM as a default approach to imputation.
Problem

Research questions and friction points this paper is trying to address.

PMM causes bias in missing-at-random data imputation
PMM fails with correlated predictors of missingness
PMM requires large samples for unbiased results
Innovation

Methods, ideas, or system contributions that make the work stand out.

PMM imputation biased under missing at random
Bias persists despite reduced variable correlation
PMM requires larger samples than other methods