🤖 AI Summary
This study addresses the high cost and prolonged duration associated with measuring primary outcomes in randomized trials by proposing a blinded adaptive sampling design. The method leverages auxiliary variables to dynamically adjust sampling probabilities while ensuring statisticians remain blinded to treatment assignments. Treatment effects are estimated by combining augmented inverse probability weighting (AIPW) with residual variance modeling, and variance loss upper bounds along with valid confidence intervals are derived based on the martingale central limit theorem. Compared with simple random sampling, the proposed approach improves efficiency by 12%–34% and reduces required sample sizes by 10%–26%. In a reanalysis of an antifungal trial, the method achieved a 39% reduction in variance, substantially enhancing statistical power.
📝 Abstract
In some randomized trials the primary outcome is costly or slow to measure, while auxiliary variables that predict it are available for everyone. The outcome can then be measured in a probability sample. We study an adaptive design in which the statistician who selects outcomes to measure is blinded to treatment assignment. Outcomes from an initial random sample are used to fit pooled models for the outcome and its residual variance. These set the sampling probabilities for the remaining participants and may be refitted as outcomes accumulate. After unblinding, arm means are estimated by augmented inverse probability weighting. The estimator is unbiased for the complete-data treatment difference for any working models, and a martingale central limit theorem gives Wald intervals under repeated updating. We bound the variance lost by estimating the sampling rule. The bound is linear in the error of the fitted residual variance, quadratic when the optimal probabilities are not truncated, and relates the initial sample size to the learning rate of the models. Relative to designs using treatment assignment, the blinded design loses a term due to unequal residual variances in the two arms and a term due to the conditional treatment effect, which is second order near the null. In simulations, coverage was near nominal. Adaptive sampling was 12% to 34% more efficient than simple random sampling and needed 10% to 26% fewer measured outcomes for the same precision. In a resampling study of an antifungal trial, adaptive sampling reduced the sampling variance by 39%.