🤖 AI Summary
This study addresses the integration of machine learning methods into survey sampling for accurate estimation of finite population parameters while preserving valid design-based statistical inference. It proposes a tailored double/debiased machine learning framework—adapted to survey data—for model-assisted estimation, item nonresponse imputation, and unit nonresponse adjustment. By combining cross-fitting, Neyman-orthogonal estimating equations, and inverse probability weighting, the approach effectively incorporates high-dimensional or nonparametric learners. The resulting estimators achieve root-n consistency and asymptotic normality, overcoming inferential challenges posed by sample dependence. The framework yields accurate estimates with desirable statistical properties in the first two settings, while also revealing limitations of doubly robust methods under unit nonresponse in official statistics applications.
📝 Abstract
This pedagogical review examines the use of machine learning methods in finite-population inference for survey sampling, with an emphasis on design-based validity and statistical inference. While flexible prediction tools offer substantial gains in estimation accuracy, they also introduce important challenges, primarily due to the dependence between the fitted predictors and the sample. We focus on settings in which such predictions enter survey estimation through model-assisted estimation, item nonresponse imputation, and unit nonresponse adjustment. For model-assisted estimation and item nonresponse, we show how cross-fitting and Neyman-orthogonal estimating equations can adapt ideas from double/debiased machine learning to survey data, allowing the use of high-dimensional or nonparametric learners while preserving root-n consistency and asymptotic normality under suitable conditions. In contrast, for unit nonresponse, standard inverse-probability weighting remains outcome-agnostic and operationally attractive, but this same feature makes doubly robust and orthogonal constructions harder to deploy in official statistics. We also briefly discuss related developments in small area estimation and probability/nonprobability data integration. Overall, the paper highlights both the promise of machine learning and the fundamental inferential challenges it raises for survey practice.