🤖 AI Summary
Traditional stratified randomized experiments fail to effectively leverage covariate information predictive of potential outcomes, resulting in suboptimal statistical efficiency. To address this, we propose an adaptive stratification design that integrates the strengths of stratification and regression adjustment: it employs batch-adaptive stratification, cross-batch rematching, and a stratified estimator to enable post-design, nonparametric covariate adjustment. This approach breaks the conventional paradigm that treats stratification and adjustment as mutually exclusive—achieving robustness to model misspecification while preserving randomization guarantees. Through simulations on synthetic data and real-world political science experiments, our method demonstrates substantial improvements in estimation accuracy and statistical efficiency, particularly when covariates exhibit strong predictive power for outcomes.
📝 Abstract
To increase statistical efficiency in a randomized experiment, researchers often use stratification (i.e., blocking) in the design stage. However, conventional practices of stratification fail to exploit valuable information about the predictive relationship between covariates and potential outcomes. In this paper, I introduce an adaptive stratification procedure for increasing statistical efficiency when some information is available about the relationship between covariates and potential outcomes. I show that, in a paired design, researchers can rematch observations across different batches. For inference, I propose a stratified estimator that allows for nonparametric covariate adjustment. I then discuss the conditions under which researchers should expect gains in efficiency from stratification. I show that stratification complements rather than substitutes for regression adjustment, insuring against adjustment error even when researchers plan to use covariate adjustment. To evaluate the performance of the method relative to common alternatives, I conduct simulations using both synthetic data and more realistic data derived from a political science experiment. Results demonstrate that the gains in precision and efficiency can be substantial.