Score
Designs, implements, and analyzes randomized controlled experiments (A/B tests) and their supporting infrastructure to assign treatments, collect and validate metrics, and measure treatment effects against baselines using appropriate statistical methods. Builds experiment frameworks and platforms, ensures validity and production-scale reliability (including assessing latency and throughput impacts), and interprets results to inform system or feature decisions.
This study addresses the lack of a systematic taxonomy in Bayesian A/B testing, which has led to the conflation of prior selection and stopping rules, resulting in methodological misuse and performance risks. The authors propose a three-tier classification framework encompassing posterior consistency, error rate control under Bayes factor–based stopping, and empirical Bayes–driven false discovery rate calibration. This work provides the first comprehensive formalization of Bayesian A/B testing methodologies, demonstrating that Bayes factor stopping is approximately optimal across a range of loss functions and establishing empirical Bayes as the sole viable route to achieving third-tier calibration. Simulations reveal that flat priors combined with posterior-based stopping amount to unprincipled peeking, that well-calibrated empirical Bayes priors substantially reduce estimation error, and that expected loss–based stopping minimizes regret only when the deployment cost of null effects is negligible.
In industrial A/B experiments, weak treatment effects often lead to insufficient statistical power. Existing methods leverage only binary trigger observations—i.e., whether outputs differ between groups—ignoring the magnitude of such differences, while full annotation of trigger intensity is prohibitively costly. This paper introduces “trigger intensity” into the A/B evaluation framework for the first time, proposing two estimation paradigms: omniscient (full knowledge) and sampling-based (partial knowledge). We theoretically prove that sampling bias asymptotically vanishes as sample size increases. Our method integrates trigger identification, stratified sampling, bias analysis, and Monte Carlo simulation. Validated on real-world business data, the omniscient approach reduces standard error by 85%, while the sampling-based approach achieves a 36.48% reduction—significantly improving estimation accuracy and statistical power.
Accurately estimating the causal effects of long-term product interventions—such as UI redesigns or recommendation algorithm updates—in digital platforms remains challenging, as conventional short-term A/B tests fail to capture delayed and evolving impacts. To address this, we propose the first causal inference framework specifically designed for estimating long-term treatment effects. Our approach disentangles time-varying confounding from lagged treatment effects by explicitly modeling treatment duration as a key covariate. It integrates structural time-series modeling, doubly robust estimation, and dynamic causal graphs to enable counterfactual effect estimation without requiring costly long-duration experiments. Evaluated on real-world platform data, our method reduces long-term effect estimation error by 42% and achieves high-fidelity predictions across core metrics—including user retention rate and click-through rate—thereby significantly improving both the reliability and efficiency of long-horizon strategy evaluation.
In A/B testing, rigorously evaluating novel estimation algorithms—when the true treatment effect is unobserved—remains a fundamental methodological challenge. This paper establishes, for the first time, a comprehensive theoretical framework for estimation and inference based on sample splitting: it derives the asymptotic distribution of sample-split estimators and characterizes their bias structure relative to full-sample performance; introduces a bias–variance trade-off analytical paradigm and proposes a correction-based confidence interval construction method. Leveraging statistical inference, asymptotic theory, Monte Carlo simulation, and empirical validation, the framework enables robust, production-grade evaluation of new algorithms within industrial A/B testing platforms. Theoretical results are thoroughly validated via simulation studies. The proposed infrastructure enhances A/B testing methodology by delivering an interpretable, reproducible, and deployable evaluation system.
In A/B testing, control variates and regression adjustment are widely used variance reduction techniques, yet their theoretical relationship remains unclear, their methodological frameworks are disjointed, and both have long been confined to design-driven paradigms. Method: This paper establishes, for the first time, a formal equivalence between these two approaches and proposes a novel grouped coefficient estimation method that unifies design-based and model-based estimation frameworks—enabling a paradigm shift from design-driven to model-driven inference. Contribution/Results: Theoretical analysis demonstrates improved estimation accuracy and statistical power. Empirical validation on millions of real-world experiments at ByteDance confirms efficacy: the proposed method has been fully deployed in its online experimentation platform, yielding an average 12.3% increase in statistical significance and a 19.6% improvement in detection sensitivity.
A/B testing is costly, necessitating effective pre-screening approaches. This work proposes a Simulated Randomized Controlled Trial (S-RCT) framework that leverages AI agents to simulate experimental outcomes based on user profiles and intervention descriptions. It introduces a novel two-stage pre-experiment calibration protocol and a within-subject design, alongside a two-level error decomposition that disentangles agent approximation error from sampling error. The framework is compatible with arbitrary behavioral models and requires no agent customization. Evaluation across 67 real-world marketing A/B tests demonstrates that post-calibration predictions reduce mean squared error by approximately 77-fold, the within-subject design lowers standard errors by 2.4-fold, and effect direction consistency reaches 0.70, substantially enhancing simulation accuracy and practical utility.
This study addresses the challenge of efficiently identifying practically meaningful treatment effects under resource constraints and concurrent experimentation, where conventional resource allocation strategies—optimized to minimize mean squared error (MSE)—often prove suboptimal. The authors propose a novel framework that shifts the objective toward minimizing the worst-case Type II error (i.e., miss rate) by leveraging statistical power. They develop a variance inflation mechanism with a correction factor, tailored to scenarios where outcome standard deviations are either known or estimated from pilot data, and formulate optimization models under three distinct risk criteria. A fully data-driven Surrogate-S algorithm is introduced to implement the approach without requiring ground-truth variance information. Theoretical analysis demonstrates the potential inefficiency of MSE-oriented strategies in detection tasks, while numerical experiments show that the proposed method achieves near-optimal performance using only pilot-based variance estimates.
This study addresses the risk of inference bias and potential failure of the CUPED method in online A/B testing under complex experimental conditions. The authors systematically investigate five critical issues related to variance reduction with CUPED, and for the first time delineate its applicability boundaries in designs such as multi-arm experiments and two-stage sampling. To overcome these limitations, they propose a robust variance estimation approach tailored to such settings. Through rigorous theoretical analysis and large-scale empirical validation, the proposed method significantly improves inference accuracy. The solution has been successfully deployed in ByteDance’s experimentation platform, demonstrating its practical effectiveness and scalability.
In A/B testing, short-term data noise, time-varying treatment effects (non-stationarity), and cross-experiment heterogeneity (e.g., product, user, and seasonal differences) lead to biased and high-variance effect estimates. To address this, we propose a local empirical Bayes framework that jointly enables temporal and contextual adaptation in control group construction—the first such approach. It dynamically selects the most relevant neighboring experiments and historical time windows via time-series alignment and context-aware similarity matching, performing localized aggregation instead of global pooling to avoid bias amplification and signal dilution. Theoretical analysis and empirical evaluation demonstrate that our method significantly reduces estimation variance while maintaining bias control, thereby improving the accuracy and stability of early-stage decisions. This yields a scalable, heterogeneity-aware modeling paradigm for high-temporal-resolution and high-reliability A/B testing.
This study addresses the persistent ambiguity in classifying repeated measures experimental designs, which often arises from conceptual confusion. To resolve this issue, the authors systematically clarify the core characteristics of such designs and propose a novel classification framework grounded in experimental units and randomization strategies. For the first time in this context, Hasse diagrams are introduced to visually represent the hierarchical structure of these designs. This approach effectively distinguishes among various types of repeated measures designs, eliminates terminological ambiguities, and substantially enhances both the rigor and interpretability of experimental planning and reporting.
Online A/B testing often necessitates interim analyses due to resource constraints; however, frequent “peeking” at accumulating data inflates Type I error rates and compromises conclusion validity. To address this, we propose a Bayesian predictive probability–based framework for safe interim evaluation. Our method is the first to enable efficient, closed-form computation of Bayesian predictive probabilities—bypassing numerical integration—thereby supporting scalable deployment and real-time experiment health monitoring. It rigorously balances statistical validity with engineering practicality. Evaluated on large-scale, real-world A/B tests from Instagram, the system significantly reduces false positive rates while ensuring reliable early stopping decisions. Deployed as a production-grade infrastructure within Meta’s experimentation platform, it enhances both experimental fidelity and resource efficiency across thousands of concurrent experiments.