🤖 AI Summary
While randomized experiments are the gold standard for causal inference, certain assignment mechanisms can induce estimation bias and misleading conclusions. Existing restricted randomization methods are limited to two-arm designs and average treatment effect (ATE) estimation, rendering them inadequate for complex settings involving interference, clustered structures, or multiple target estimands. This paper introduces the *Inference-Guided Randomization* (IGR) framework: a pre-registered, domain-knowledge-informed approach that dynamically selects high-quality assignments via design criteria; the first to extend restricted randomization to diverse estimands—including cluster-level effects, direct and indirect effects—and nonstandard experimental structures; and integrating simulation-based filtering, Monte Carlo evaluation, and reproducible design workflows. In simulations of behavioral health and educational interventions, IGR substantially improves estimation accuracy and statistical power while ensuring design transparency, resistance to p-hacking, and practical feasibility.
📝 Abstract
Randomized experiments are considered the gold standard for estimating causal effects. However, out of the set of possible randomized assignments, some may be more likely to produce poor effect estimates and misleading conclusions. Restricted randomization is an experimental design strategy that filters out undesirable treatment assignments, but its application has primarily been limited to ensuring covariate balance in two-arm studies where the target estimand is the average treatment effect. Other experimental settings with different design desiderata and target effect estimands could also stand to benefit from a restricted randomization approach. We introduce inspection-guided randomization (IGR), a transparent and flexible framework for restricted randomization that filters out undesirable treatment assignments by inspecting assignments against analyst-specified, domain-informed design desiderata. In IGR, the acceptable treatment assignments are locked in ex ante and preregistered in the trial protocol, thus safeguarding against p-hacking and promoting reproducibility. Through illustrative simulation studies motivated by behavioral health and education interventions, we demonstrate how IGR can improve effect estimates compared to benchmark designs in experiments involving interference and group formation.