π€ AI Summary
This study addresses the severe undercoverage of confidence intervals when estimating population average treatment effects using flexible machine learning models under complex survey designs, such as stratified multistage sampling. The authors propose a survey-weighted targeted maximum likelihood estimator (TMLE) with cross-fitting implemented at the primary sampling unit (PSU) level, combined with Taylor linearization of the influence function to obtain design-consistent variance estimates. Theory and simulations demonstrate for the first time that valid inference requires cross-fitting specifically at the PSU levelβinternal cross-validation is insufficient. In NHANES-like simulations, the proposed method achieves stable coverage of 93%β95%, substantially outperforming single-fit TMLE (as low as 22%) and internal cross-validation (85%β88%). Empirical analyses of four NHANES datasets confirm its practical utility, and accompanying open-source software has been released.
π Abstract
Cross-fitting is not a refinement of survey-weighted causal machine learning but, once the nuisances are flexible, what restores valid inference. We study the population average treatment effect under a stratified multistage design, estimated by a survey-aware targeted maximum likelihood estimator (TMLE) whose variance is obtained by Taylor-series linearization of the influence function, treating the primary sampling unit as the replication unit. Our central result, established in theory and simulation, is that this validity turns on cross-fitting at the cluster level. Once flexible learners cross a complexity (Donsker) boundary, single-fit survey TMLE can severely under-cover, and internal cluster-aware cross-validation does not substitute for cross-fitting; among the estimators we evaluate, only out-of-fold fitting at the cluster level restores valid coverage. In simulations spanning a many-PSU and an NHANES-like design, on a diverse ensemble the single-fit and internal cross-validation estimators cover at about 0.89-0.91 and 0.85-0.88 while the cross-fitted estimator holds at 0.93-0.95, and an aggressively grown learner drives single-fit coverage to 0.22. Two scope choices are deliberate: survey-weighted point estimation is prior work, and the nuisance product-rate condition is assumed and probed empirically. Within these conditions we prove asymptotic normality and design-consistency of the linearization variance. Four NHANES analyses and open-source software illustrate the method.