run a/b experiments

Designs, implements, and analyzes randomized controlled experiments (A/B tests) and their supporting infrastructure to assign treatments, collect and validate metrics, and measure treatment effects against baselines using appropriate statistical methods. Builds experiment frameworks and platforms, ensures validity and production-scale reliability (including assessing latency and throughput impacts), and interprets results to inform system or feature decisions.

runabexperiments

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.9
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

In industrial A/B experiments, weak treatment effects often lead to insufficient statistical power. Existing methods leverage only binary trigger observations—i.e., whether outputs differ between groups—ignoring the magnitude of such differences, while full annotation of trigger intensity is prohibitively costly. This paper introduces “trigger intensity” into the A/B evaluation framework for the first time, proposing two estimation paradigms: omniscient (full knowledge) and sampling-based (partial knowledge). We theoretically prove that sampling bias asymptotically vanishes as sample size increases. Our method integrates trigger identification, stratified sampling, bias analysis, and Monte Carlo simulation. Validated on real-world business data, the omniscient approach reduces standard error by 85%, while the sampling-based approach achieves a 36.48% reduction—significantly improving estimation accuracy and statistical power.

Analyzing bias in sampling-based evaluation methodsEnhancing A/B experiment precision via trigger intensityReducing cost of trigger observation detection

Estimating Effects of Long-Term Treatments

Jul 07, 2023
SH
Shan Huang
🏛️ The University of Hong Kong | University of California, Davis | Boston University | Tencent, Inc.

Accurately estimating the causal effects of long-term product interventions—such as UI redesigns or recommendation algorithm updates—in digital platforms remains challenging, as conventional short-term A/B tests fail to capture delayed and evolving impacts. To address this, we propose the first causal inference framework specifically designed for estimating long-term treatment effects. Our approach disentangles time-varying confounding from lagged treatment effects by explicitly modeling treatment duration as a key covariate. It integrates structural time-series modeling, doubly robust estimation, and dynamic causal graphs to enable counterfactual effect estimation without requiring costly long-duration experiments. Evaluated on real-world platform data, our method reduces long-term effect estimation error by 42% and achieves high-fidelity predictions across core metrics—including user retention rate and click-through rate—thereby significantly improving both the reliability and efficiency of long-horizon strategy evaluation.

Decomposing long-term effects using user attributes and short-term metricsEstimating long-term treatment effects from short-term A/B testing dataValidating framework with large-scale real-world experiments on WeChat

In A/B testing, rigorously evaluating novel estimation algorithms—when the true treatment effect is unobserved—remains a fundamental methodological challenge. This paper establishes, for the first time, a comprehensive theoretical framework for estimation and inference based on sample splitting: it derives the asymptotic distribution of sample-split estimators and characterizes their bias structure relative to full-sample performance; introduces a bias–variance trade-off analytical paradigm and proposes a correction-based confidence interval construction method. Leveraging statistical inference, asymptotic theory, Monte Carlo simulation, and empirical validation, the framework enables robust, production-grade evaluation of new algorithms within industrial A/B testing platforms. Theoretical results are thoroughly validated via simulation studies. The proposed infrastructure enhances A/B testing methodology by delivering an interpretable, reproducible, and deployable evaluation system.

Derives asymptotic distributions and constructs valid confidence intervalsDevelops a theoretical framework for sample splitting in A/B testingValidates results through simulations and provides implementation guidance

In A/B testing, control variates and regression adjustment are widely used variance reduction techniques, yet their theoretical relationship remains unclear, their methodological frameworks are disjointed, and both have long been confined to design-driven paradigms. Method: This paper establishes, for the first time, a formal equivalence between these two approaches and proposes a novel grouped coefficient estimation method that unifies design-based and model-based estimation frameworks—enabling a paradigm shift from design-driven to model-driven inference. Contribution/Results: Theoretical analysis demonstrates improved estimation accuracy and statistical power. Empirical validation on millions of real-world experiments at ByteDance confirms efficacy: the proposed method has been fully deployed in its online experimentation platform, yielding an average 12.3% increase in statistical significance and a 19.6% improvement in detection sensitivity.

Analyzing statistical properties and theoretical connections between frameworksBridging control variates and regression adjustment methodsProviding guidance for variance reduction in A/B testing

A/B testing is costly, necessitating effective pre-screening approaches. This work proposes a Simulated Randomized Controlled Trial (S-RCT) framework that leverages AI agents to simulate experimental outcomes based on user profiles and intervention descriptions. It introduces a novel two-stage pre-experiment calibration protocol and a within-subject design, alongside a two-level error decomposition that disentangles agent approximation error from sampling error. The framework is compatible with arbitrary behavioral models and requires no agent customization. Evaluation across 67 real-world marketing A/B tests demonstrates that post-calibration predictions reduce mean squared error by approximately 77-fold, the within-subject design lowers standard errors by 2.4-fold, and effect direction consistency reaches 0.70, substantially enhancing simulation accuracy and practical utility.

A/B testingAI agentsexperimentation

Latest Papers

What's happening recently
View more

This study addresses the challenge of efficiently identifying practically meaningful treatment effects under resource constraints and concurrent experimentation, where conventional resource allocation strategies—optimized to minimize mean squared error (MSE)—often prove suboptimal. The authors propose a novel framework that shifts the objective toward minimizing the worst-case Type II error (i.e., miss rate) by leveraging statistical power. They develop a variance inflation mechanism with a correction factor, tailored to scenarios where outcome standard deviations are either known or estimated from pilot data, and formulate optimization models under three distinct risk criteria. A fully data-driven Surrogate-S algorithm is introduced to implement the approach without requiring ground-truth variance information. Theoretical analysis demonstrates the potential inefficiency of MSE-oriented strategies in detection tasks, while numerical experiments show that the proposed method achieves near-optimal performance using only pilot-based variance estimates.

A/B testingexperiment-rich regimeresource allocation

This study addresses the risk of inference bias and potential failure of the CUPED method in online A/B testing under complex experimental conditions. The authors systematically investigate five critical issues related to variance reduction with CUPED, and for the first time delineate its applicability boundaries in designs such as multi-arm experiments and two-stage sampling. To overcome these limitations, they propose a robust variance estimation approach tailored to such settings. Through rigorous theoretical analysis and large-scale empirical validation, the proposed method significantly improves inference accuracy. The solution has been successfully deployed in ByteDance’s experimentation platform, demonstrating its practical effectiveness and scalability.

A/B testingCUPEDmulti-arm experiments

Learning Across Experiments and Time: Tackling Heterogeneity in A/B Testing

Nov 26, 2025
XL
Xinran Li
🏛️ University of Science and Technology of China

In A/B testing, short-term data noise, time-varying treatment effects (non-stationarity), and cross-experiment heterogeneity (e.g., product, user, and seasonal differences) lead to biased and high-variance effect estimates. To address this, we propose a local empirical Bayes framework that jointly enables temporal and contextual adaptation in control group construction—the first such approach. It dynamically selects the most relevant neighboring experiments and historical time windows via time-series alignment and context-aware similarity matching, performing localized aggregation instead of global pooling to avoid bias amplification and signal dilution. Theoretical analysis and empirical evaluation demonstrate that our method significantly reduces estimation variance while maintaining bias control, thereby improving the accuracy and stability of early-stage decisions. This yields a scalable, heterogeneity-aware modeling paradigm for high-temporal-resolution and high-reliability A/B testing.

Addresses noisy treatment effect estimates in A/B testing due to short horizonsHandles cross-experiment variability in product, users, and seasonalitySolves bias from pooling experiments with temporal evolution and heterogeneity

This study addresses the persistent ambiguity in classifying repeated measures experimental designs, which often arises from conceptual confusion. To resolve this issue, the authors systematically clarify the core characteristics of such designs and propose a novel classification framework grounded in experimental units and randomization strategies. For the first time in this context, Hasse diagrams are introduced to visually represent the hierarchical structure of these designs. This approach effectively distinguishes among various types of repeated measures designs, eliminates terminological ambiguities, and substantially enhances both the rigor and interpretability of experimental planning and reporting.

experimental designexperimental unitsHasse diagrams

Bayesian Predictive Probabilities for Online Experimentation

Nov 09, 2025
AZ
Abbas Zaidi
🏛️ Meta Inc

Online A/B testing often necessitates interim analyses due to resource constraints; however, frequent “peeking” at accumulating data inflates Type I error rates and compromises conclusion validity. To address this, we propose a Bayesian predictive probability–based framework for safe interim evaluation. Our method is the first to enable efficient, closed-form computation of Bayesian predictive probabilities—bypassing numerical integration—thereby supporting scalable deployment and real-time experiment health monitoring. It rigorously balances statistical validity with engineering practicality. Evaluated on large-scale, real-world A/B tests from Instagram, the system significantly reduces false positive rates while ensuring reliable early stopping decisions. Deployed as a production-grade infrastructure within Meta’s experimentation platform, it enhances both experimental fidelity and resource efficiency across thousands of concurrent experiments.

Addresses capacity constraints in online A/B testing requiring interim analysesEnables reliable interim decisions without compromising experiment fidelitySolves error-prone peeking practices that inflate type-I error rates

Hot Scholars

KG

Kun Gai

Senior Director & Researcher, Alibaba Group
Machine LearningComputational Advertising
LH

Lantao Hu

Kuaishou Inc.
data miningrecommeder system
GZ

Guorui Zhou

Unknown affiliation
Recommender System,Advertising,Artificial Intelligence,Machine Learning,NLP
JC

Jiangxia Cao

Kuaishou Tech
RecSysLow-Resource Large Model