๐ค AI Summary
In A/B testing, short-term data noise, time-varying treatment effects (non-stationarity), and cross-experiment heterogeneity (e.g., product, user, and seasonal differences) lead to biased and high-variance effect estimates. To address this, we propose a local empirical Bayes framework that jointly enables temporal and contextual adaptation in control group constructionโthe first such approach. It dynamically selects the most relevant neighboring experiments and historical time windows via time-series alignment and context-aware similarity matching, performing localized aggregation instead of global pooling to avoid bias amplification and signal dilution. Theoretical analysis and empirical evaluation demonstrate that our method significantly reduces estimation variance while maintaining bias control, thereby improving the accuracy and stability of early-stage decisions. This yields a scalable, heterogeneity-aware modeling paradigm for high-temporal-resolution and high-reliability A/B testing.
๐ Abstract
A/B testing plays a central role in data-driven product development, guiding launch decisions for new features and designs. However, treatment effect estimates are often noisy due to short horizons, early stopping, and slowly accumulating long-tail metrics, making early conclusions unreliable. A natural remedy is to pool information across related experiments, but naive pooling potentially fails: within experiments, treatment effects may evolve over time, so mixing early and late outcomes without accounting for nonstationarity induces bias; across experiments, heterogeneity in product, user population, or season dilutes the signal with unrelated noise. These issues highlight the need for pooling strategies that adapt to both temporal evolution and cross-experiment variability. To address these challenges, we propose a local empirical Bayes framework that adapts to both temporal and cross-experiment heterogeneity. Throughout an experiment's timeline, our method builds a tailored comparison set: time-aware within the experiment to respect nonstationarity, and context-aware across experiments to draw only from comparable counterparts. The estimator then borrows strength selectively from this set, producing stabilized treatment effect estimates that remain sensitive to both time dynamics and experimental context. Through theoretical analysis and empirical evaluation, we show that the proposed local pooling strategy consistently outperforms global pooling by reducing variance while avoiding bias. Our proposed framework enhances the reliability of A/B testing under practical constraints, thereby enabling more timely and informed decision-making.