🤖 AI Summary
This study addresses the challenge in randomized experiments where conventional sample allocation methods fail to simultaneously account for deployment relevance and statistical precision for a target population. The authors propose TWNA, a two-stage stratified design: an initial pilot stage estimates stratum-specific treatment effect variances, which then inform a joint optimization of final-stage sample sizes and treatment probabilities to enhance estimation precision of the target-weighted group average treatment effect (GATE). TWNA is the first method to unify deployment weights and statistical difficulty within a single optimization framework, yielding a closed-form optimal allocation rule. It further extends robustly to complex settings involving uncertainty in target population composition, skewed outcomes, or rare events. Simulations and empirical analyses demonstrate that TWNA substantially improves estimation accuracy and resource efficiency, particularly for critical yet hard-to-estimate subgroups.
📝 Abstract
Randomized experiments are often run in one population to guide decisions in another. Allocating by experimental proportions wastes budget on groups that rarely appear in deployment, whereas allocating by deployment proportions under-samples groups that are hard to measure precisely. We propose \textbf{TWNA} (Target-Weighted Neyman Allocation), a two-stage stratified design that uses pilot estimates of group--arm outcome variances to allocate final-stage sample sizes and treatment probabilities for target-weighted group average treatment effect (GATE) precision. The oracle rule has a closed form and balances deployment importance with statistical difficulty; the plug-in rule recovers it as pilot variance estimates stabilize. We also extend TWNA to handle uncertainty about deployment composition, remaining robust whether the target mix is roughly known or entirely unknown. Finally, we distinguish this weight robustness from a pilot-robust variant for skewed, rare-event, or contaminated outcomes. Simulations and real-covariate benchmarks show the largest gains when groups are both deployment-important and difficult to measure.