design randomized experiments

Designs randomized assignment and sampling schemes, randomized perturbation and response mechanisms, and randomized algorithmic components (e.g., randomized sketching, randomized SVD, randomized rounding) to implement experiments and scalable data collection. Builds protocols for calibration and pre-registration, implements privacy-preserving randomized reporting, and analyzes stochastic variability to estimate causal effects, conduct randomized testing, and produce valid inference from randomized field trials.

designrandomizedexperiments

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.73
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$194K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Selective Randomization Inference for Adaptive Experiments

May 11, 2024
TF
Tobias Freidling
🏛️ École polytechnique fédérale de Lausanne | University of Cambridge | University of Southern California

In adaptive experiments, data-driven design adjustments invalidate conventional statistical inference; existing methods suffer from narrow applicability and strong assumptions. This paper proposes a selective randomization inference framework: it models the data-generating process via a directed acyclic graph (DAG) and implements conditional-on-selection inference within randomization tests. It is the first systematic integration of this principle into randomization-based inference—requiring neither i.i.d. nor parametric modeling assumptions, and accommodating arbitrary adaptive experimental designs. To address disconnected confidence intervals, we innovatively introduce a holdout-unit method. Theoretically and empirically, our approach strictly controls selective Type-I error and constructs valid confidence intervals for homogeneous treatment effects. It substantially outperforms conventional methods in both robustness and generality.

Addresses data-dependent null hypotheses and experimental designsControls selective type-I error without modeling assumptionsDevelops selective randomization inference for adaptive experiments

Identification and Inference on Treatment Effects under Covariate-Adaptive Randomization and Imperfect Compliance

Jun 12, 2024
FA
Federico A. Bugni
🏛️ Northwestern University | UC Berkeley

This paper addresses identification and inference for the average treatment effect (ATE) and average treatment effect on the treated (ATT) under covariate-adaptive randomization (CAR) with noncompliance. First, it precisely characterizes the sharp identification sets for ATE and ATT within the CAR framework—where sampling yields non-i.i.d. data—using extremal identification theory. Second, it proposes boundary estimators and confidence interval constructions that are both consistent and asymptotically efficient. Third, it develops a hybrid weighting strategy that jointly leverages empirical sampling frequencies and known randomization probabilities; it establishes that ATE estimation achieves asymptotic efficiency using only empirical frequencies, whereas ATT estimation requires the numerator to employ true compliance probabilities and the denominator empirical frequencies. These results provide both theoretical foundations and practical guidelines for robust causal inference in CAR-based randomized controlled trials.

Addresses imperfect compliance in randomized controlled trialsIdentifies treatment effects under covariate-adaptive randomizationProvides inference methods for ATE and ATT bounds

Multiple Randomization Designs

Dec 27, 2021
LM
Lorenzo Masoero
🏛️ Amazon | University of Washington | Stanford University

Traditional randomized controlled trials (RCTs) suffer from cross-group spillovers and interference in multi-population interaction settings—e.g., buyer-seller or creator-subscriber systems—leading to biased estimation of the average treatment effect (ATE). This paper introduces the first systematic multidimensional randomization design framework, relaxing the conventional single-layer randomization assumption. By integrating hierarchical randomization, cross-group assignment, and potential outcomes modeling, it jointly identifies both the ATE and cross-group interference effects. We establish theoretical guarantees: the proposed unbiased estimator is consistent and asymptotically normal. Simulation results demonstrate substantially higher statistical power compared to standard RCTs. This work extends the scope of causal questions addressable through experimental design and provides a rigorous foundation for causal inference in complex intervention environments—particularly platform economies—where interdependent user behaviors induce non-negligible interference.

Addressing interference effects in multi-population experimental settingsDeveloping statistical methods for analyzing multiple randomization designsProposing new designs for experiments with interacting populations

The Optimality of Blocking Designs in Equally and Unequally Allocated Randomized Experiments with General Response

Dec 04, 2022
DA
David Azriel
🏛️ The Technion | The Wharton School of the University of Pennsylvania | Queens College, CUNY

This paper investigates the performance of the difference-in-means estimator in two-arm randomized experiments under continuous, binary, proportion, and survival endpoints, considering both equal and unequal allocation within Neyman and superpopulation modeling frameworks. Methodologically, it integrates randomization inference, minimax analysis, asymptotic statistics, and NP-hardness arguments, complemented by Monte Carlo simulations. The study establishes, for the first time, that Fisher’s blocked design is asymptotically optimal under Kapelner’s tail criterion. It systematically characterizes the theoretical boundaries between complete randomization (Neyman model) and deterministic perfect balance (superpopulation model), identifying blocked design as the optimal compromise. Both theoretical analysis and simulation results consistently demonstrate that blocking substantially reduces estimation error and robustly outperforms complete randomization and pairwise matching across all endpoint types and allocation ratios.

Compares optimal designs under Neyman and population randomization modelsEvaluates difference-in-means estimator performance in two-arm experimentsProves blocking design achieves asymptotically optimal experimental randomness

Design-based Estimation Theory for Complex Experiments

Nov 12, 2023
HC
Haoge Chang
🏛️ Columbia University

This paper addresses complex randomized experiments subject to interference between units—such as social network interventions—where standard causal inference assumptions fail. Method: We develop a design-based theoretical framework for estimating treatment effects, introducing a family of design-compatible estimators and a scalar, interpretable measure of “experimental complexity.” We establish its theoretical connection to design variance, derive the asymptotic variance lower bound for unbiased estimation under arbitrary designs, and propose a consistent variance estimator. Contributions/Results: Through interference modeling, design-based inference foundations, and network experiment simulations, we validate our approach on real-world social network data from an insurance adoption study. Our estimators achieve significantly improved estimation accuracy and consistent variance estimation compared to existing methods, providing a theoretically rigorous yet practically implementable analytical framework for complex experimental designs.

Developing design-based estimation theory for arbitrary designsEstimating treatment effects in complex randomized experimentsProposing new estimators with favorable asymptotic properties

Latest Papers

What's happening recently
View more

This work addresses the challenge of finite-sample inference for individual regression coefficients in fixed-design linear models when errors exhibit dependence or heteroskedasticity. The authors propose a unified randomization testing framework based on group permutations, which rigorously controls Type I error under exchangeable errors and enhances power through design-dependent geometric separation. The approach is further extended to non-exchangeable settings, establishing quantitative robustness for approximately symmetric errors. The study proves that the resulting Type I error bound of level $2\alpha$ is tight and, by integrating a constructive algorithm, sub-Gaussian analysis, and conformal inference, achieves substantially improved power under heavy-tailed designs while preserving finite-sample validity.

exchangeabilityfinite-sample inferencelinear model

This work addresses a key limitation of the standard debiasing approach for the Randomized Response (RR) mechanism under local differential privacy, which often yields negative estimates that lack interpretability as histogram probabilities. The paper presents the first closed-form maximum likelihood estimator (MLE) for the true underlying distribution under RR, circumventing the high computational cost of iterative Bayesian updating (IBU) algorithms. The proposed estimator is theoretically elegant, computationally efficient, and guarantees non-negative, accurate reconstruction of the original histogram. Empirical evaluations demonstrate that this analytical solution significantly outperforms existing correction methods in both estimation accuracy and computational efficiency, offering practitioners a reliable and scalable tool for privacy-preserving frequency estimation.

Histogram EstimationLocal Differential PrivacyMaximum Likelihood Estimate

Classical randomized experiments struggle to identify causal effects in market platforms featuring cross-group strategic interactions and complex spillovers. To address this, we propose a novel multi-stage randomization design and the first finite-sample valid inferential framework that explicitly accounts for interference. We define composite causal parameters—including average direct, primary, and multiple types of spillover effects—that are both identifiable and substantively interpretable. Our method integrates graph-based randomization, hierarchical–clustered randomization, inverse-probability weighting, and Hájek-type bias correction, and establishes a finite-sample central limit theorem. We prove that all estimators achieve √n-consistency and asymptotic normality, ensuring statistical validity while substantially improving estimation precision for spillover effects. The framework is scalable and directly applicable to large-scale online marketplace experiments.

Designs address interference in strategic agent interactionsEstimators capture complex spillover effects in marketplacesMethods derive properties for direct and spillover effects

Design-based finite-sample analysis for regression adjustment

Nov 19, 2025
DS
Dogyoon Song
🏛️ University of California, Davis

This paper addresses finite-sample inference for regression-adjusted average treatment effect (ATE) estimation in high-dimensional randomized experiments where $p > n$. We propose a design-based non-asymptotic analytical framework—distinct from existing approaches relying on correct model specification or large-sample approximations. For the first time, we establish a non-asymptotic theory for high-dimensional regression-adjusted estimators, leveraging Stein’s exchangeable pair technique, Doob martingales, and Freedman’s inequality to characterize how the geometric structure of covariates governs estimation precision. The resulting confidence intervals are instance-adaptive and data-driven, with explicitly computable widths that remain informative even when $p > n$. This work provides the first rigorous, practical, design-based theoretical foundation for robust causal inference in small-sample, high-dimensional settings.

Analyzing regression adjustment for ATE estimation in high-dimensional randomized experimentsDeveloping finite-sample confidence intervals when covariates exceed observationsInvestigating covariate geometry's impact on concentration and bias properties

This study addresses the challenge of unbiased comparison between two pre-specified matching mechanisms in finite populations, where interference from one mechanism can contaminate the evaluation of the other. The authors propose an alternating-path randomization design based on a decomposition of the disagreement set, which requires no assumptions about outcome or behavioral models. By leveraging unique alternating paths and cycle decompositions from graph theory—such as augmenting paths and Eulerian cycle decompositions—the method constructs a controlled interference structure and sets the optimal randomization probability to √2−1 to minimize worst-case variance. The approach integrates Horvitz–Thompson estimation with minimax optimization, establishing unbiasedness of the estimator, proving a finite-population central limit theorem applicable to complex path structures, and extending successfully to many-to-one matching settings with capacity constraints.

disagreement setexperimental designmatching interference

Hot Scholars

RR

Ronitt Rubinfeld

Professor of Computer Science, MIT and Tel Aviv University
Computer Science: theory of computationalgorithms
MC

Maxim Chupilkin

EBRD, University of Oxford
international political economygeoeconomicseconomic statecraft
SP

Snigdha Panigrahi

University of Michigan
Selective InferenceCausal InferenceRandomizationMachine Learning
MR

Mohammad Roghani

PhD student, Stanford University
AlgorithmsTheoretical Computer Science
SS

Sandeep Silwal

Assistant Professor, University of Wisconsin-Madison
Theoretical Computer ScienceAlgorithmsMachine Learning