🤖 AI Summary
This study addresses the challenge of achieving both efficiency and robustness in parameter estimation with partially observed data under the missing-at-random (MAR) mechanism. The authors propose a unified inference framework based on generalized entropy calibration weighting. By solving a convex entropy minimization problem subject to balancing and debiasing constraints, the method constructs weights that integrate data-adaptive calibration functions, flexible machine learning predictors, cross-fitting, and propensity score modeling. This approach subsumes inverse probability weighting (IPW) and augmented IPW (AIPW) as special cases, enjoys double robustness, and attains the semiparametric efficiency bound when both outcome and propensity models are correctly specified. Simulation and empirical analyses demonstrate that the proposed estimator outperforms existing methods in terms of statistical efficiency and numerical stability, particularly exhibiting substantial gains over standard AIPW when the outcome model is misspecified.
📝 Abstract
Missing data is an universal problem in statistics. We develop a unified framework for estimating parameters defined by general estimating equations under a missing-at-random (MAR) mechanism, based on generalized entropy calibration weighting. We construct weights by minimizing a convex entropy subject to (i) balancing constraints on a data-adaptive calibration function, estimated using flexible machine-learning predictors with cross-fitting, and (ii) a debiasing constraint involving the fitted propensity score (PS) model. The resulting estimator is doubly robust, remaining consistent if either the outcome regression (OR) or the PS model is correctly specified, and attains the semiparametric efficiency bound when both models are correctly specified. Our formulation encompasses classical inverse probability weighting (IPW) and augmented IPW (AIPW) as special cases and accommodates a broad class of entropy functions. We illustrate the versatility of the approach in three important settings: semi-supervised learning with unlabeled outcomes, regression analysis with missing covariates, and causal effect estimation in observational studies. Extensive simulation studies and real-data applications demonstrate that the proposed estimators achieve greater efficiency and numerical stability than existing methods. In particular, the proposed estimator outperforms the classical AIPW estimator under the OR model misspecification.