๐ค AI Summary
Accurate estimation and inference for the average treatment effect (ATE) in multi-covariate randomized controlled trials (RCTs) remain challenging, particularly under high-dimensional covariates and small sample sizes.
Method: Building on the Neyman finite-population framework, we propose a bias-corrected regression-adjustment estimator with cross-fitting and introduce, for the first time, an HC3-type heteroskedasticity-robust standard error tailored to high-dimensional settings. We rigorously derive first- and second-order stochastic expansions of the random component of regression-adjustment estimators, identifying the source of higher-order bias in conventional inference; leveraging this insight, we design a cross-fitting procedure to eliminate bias and extend HC3 standard errors to stratified experimental designs.
Results: Simulations and reanalysis of Angrist et al. (2009)โs education RCT demonstrate that our method substantially improves estimation accuracy, confidence interval coverage, and statistical powerโespecially in small-sample and high-dimensional scenarios.
๐ Abstract
This paper is concerned with estimation and inference on average treatment effects in randomized controlled trials when researchers observe potentially many covariates. By employing Neyman's (1923) finite population perspective, we propose a bias-corrected regression adjustment estimator using cross-fitting, and show that the proposed estimator has favorable properties over existing alternatives. For inference, we derive the first and second order terms in the stochastic component of the regression adjustment estimators, study higher order properties of the existing inference methods, and propose a bias-corrected version of the HC3 standard error. The proposed methods readily extend to stratified experiments with large strata. Simulation studies show our cross-fitted estimator, combined with the bias-corrected HC3, delivers precise point estimates and robust size controls over a wide range of DGPs. To illustrate, the proposed methods are applied to real dataset on randomized experiments of incentives and services for college achievement following Angrist, Lang, and Oreopoulos (2009).