🤖 AI Summary
This study addresses the limitations of existing mean-difference estimators in randomized controlled trials, which, while unbiased, often exhibit limited precision and lack a systematic framework for selection tailored to distinct analytical objectives—such as statistical inference versus decision-making. To overcome this, the authors propose a general sample-splitting framework that reframes estimator selection as a goal-oriented comparative task, thereby avoiding the validity risks associated with post hoc choices. By estimating the distributions of metrics like mean squared error and regret across multiple estimation methods—including covariate adjustment, weighted least squares, and simple difference-in-means—the framework enables objective evaluation. Empirical analyses on Amazon supply-chain experiments and the Strengthening Democracy Challenge (25 interventions) reveal that weighted least squares performs best for inference, whereas difference-in-means yields the lowest regret in decision contexts, demonstrating for the first time that the optimal estimator varies systematically with the analytical goal.
📝 Abstract
Randomized controlled trials (RCTs) are fundamental tools for causal inference across technology companies, pharmaceutical research, and federal agencies. While the standard difference-in-means estimator provides unbiased treatment effect estimates, it often lacks precision, particularly when treatment effects are heterogeneous or outcomes exhibit heavy-tailed distributions. Although numerous precision-enhancing methods exist---from covariate adjustment techniques to variance reduction strategies---recent research demonstrates that no single estimator performs optimally across all datasets. Rather than seeking the best estimator for individual RCTs, which risks compromising scientific validity through convenient selection, we propose a principled framework for identifying optimal estimators within families of RCTs based on specific analytical goals. Our approach uses sample splitting to estimate the distribution of evaluation metrics (e.g., mean squared error, regret) across RCT families, enabling systematic comparisons between estimators while maintaining asymptotic guarantees. We demonstrate this framework using a sample of Amazon's Supply Chain Optimization Technology trials and the Strengthening Democracy Challenge dataset (25 interventions). Results reveal that optimal estimators vary significantly by analytical objective: weighted least squares performs best for inference goals, while difference-in-means minimizes regret for decision-making contexts. This work provides actionable guidance for estimator selection while preserving methodological rigor across diverse research applications.