🤖 AI Summary
This study addresses the challenge of simultaneously achieving early stopping, false positive control, and precise effect estimation in online multi-factor experiments. It proposes a framework that jointly optimizes allocation and evidence rules. Specifically, block orthogonal designs are employed to eliminate factor interference, while isotropic allocation combined with radial e-value merging attains minimax rate optimality under alternative hypotheses with unknown directions. Furthermore, by integrating arbitrary valid e-processes, maximum likelihood estimation, and batch replication techniques, the method achieves dual optimality for both testing and estimation. Simulation studies and experiments involving large language models demonstrate that this approach significantly improves stopping efficiency and estimation accuracy.
📝 Abstract
Online multifactorial experiments requires early stopping for practically negligible effects while controlling false stopping under continuous monitoring and retaining precise treatment-effect estimation. We develop a design-driven framework that chooses allocation and the evidence rule jointly, treating allocation as evidence collection and testing and estimation as complementary uses of the same evidence. For multifactor experiments with nuisance block effects and treatment-by-block interactions, we show that block-orthogonal designs remove nuisance contamination from the treatment score and form a complete class for worst-case evidence growth. Under a fixed information budget, an isotropic allocation and an explicit radial e-value jointly attain the minimax rate optimal for directionally unknown alternatives. The same allocation is universally optimal for maximum likelihood estimation of treatment effects, simultaneously achieving A-, D-, and E-optimality and establishing double optimality for testing and estimation. Batchwise replication yields an anytime-valid e-process, retains minimax rate optimality for cumulative conditional worst-case evidence growth, and preserves universal estimation optimality for the prespecified complete experiment. Simulations and a large language model prompt experiment illustrate gains in stopping efficiency and estimation precision.