run controlled experiments

Design and conduct experiments that manipulate one or more independent variables while holding or counterbalancing other factors so as to isolate and measure causal effects precisely. This includes creating randomized assignment and presentation schemes, constructing ablation or controlled-variable comparison conditions, controlling confounds, and specifying statistical analyses to detect and attribute performance or order effects.

runcontrolledexperiments

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$205K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Learning control variables and instruments for causal analysis in observational data

Jul 05, 2024
NA
Nicolas Apfel
🏛️ University of Innsbruck | University of York | University of Fribourg | Heinrich Heine University Düsseldorf

Estimating causal effects from observational data requires selecting appropriate control and instrumental variables that satisfy causal identification conditions—a challenging task often reliant on strong domain knowledge or ad hoc assumptions. Method: This paper proposes the first end-to-end joint learning framework that automatically identifies valid combinations of control and instrumental variables. Grounded in conditional independence testing, the method integrates nonparametric dependence measures with structural search optimization, ensuring statistical consistency in variable selection under mild regularity conditions. Contribution/Results: Unlike conventional approaches requiring prespecified variable sets or strong prior assumptions, our framework is fully data-driven. In simulations, it achieves significantly higher variable identification accuracy. Empirically, applied to the Job Corps study, its estimated treatment effect closely aligns with results from the randomized controlled trial—demonstrating both validity and robustness in real-world causal inference.

Detects control variables and instruments for causal analysis in observational dataLearns partition of instruments and control variables from observed dataTests joint existence of instruments and control variables using machine learning

Synthetic Controls for Experimental Design

Aug 04, 2021
AA
Alberto Abadie
🏛️ MIT | Boston University

In large-scale aggregate-unit experiments (e.g., markets), conventional randomized treatment assignment often yields severe baseline imbalance due to extremely few treated units, leading to biased causal estimates. To address this, we systematically integrate the synthetic control method into experimental design, proposing a non-randomized treatment allocation mechanism: dynamically constructing a weighted synthetic control group based on pre-treatment covariates. We further develop配套 components—including counterfactual prediction, distance-driven unit matching, robust variance estimation, and a novel confidence interval construction procedure. Theoretically, our estimator is proven consistent and asymptotically normal. Empirically, it reduces estimation bias by 40–65% relative to standard randomization and substantially improves statistical power. Our core contribution is a new causal inference paradigm for small-N aggregate experiments—rigorous in inference, unbiased under mild assumptions, and highly interpretable.

Addresses experimental design for large aggregate unitsProposes synthetic control designs for accurate estimationReduces bias in treated and control group selection

Causal Inference in Counterbalanced Within-Subjects Designs

May 06, 2025
JH
Justin Ho
🏛️ Harvard University | University of California, Berkeley

Counterbalanced within-subject experiments risk invalid causal inference due to unverifiable and often violated assumptions—particularly the symmetry and cancelability of carryover effects. Method: We introduce “sequential exchangeability” as a formal identification assumption within the potential outcomes framework, rigorously exposing inherent limitations of counterbalancing; we then develop actionable strategies—including diagnostic tests, optimized washout periods, covariate adjustment, and alternative sequence designs—grounded in causal identification theory, sequential randomization modeling, and sensitivity analysis. Contribution/Results: Our work delineates precise validity boundaries for counterbalanced designs, providing rigorous, practical guidelines for within-subject experimentation in psychology, human-computer interaction, and related fields. By addressing foundational identifiability concerns, it substantially enhances the reliability of causal inference in repeated-measures settings.

Challenges in causal inference with counterbalanced within-subjects designsNeed for alternative methods to ensure valid causal inferenceUnverifiable assumptions about symmetric carryover effects in counterbalancing

Combining an experimental study with external data: study designs and identification strategies

Jun 05, 2024
LU
Lawson Ung
🏛️ Harvard T.H. Chan School of Public Health | Dartmouth Geisel School of Medicine | Beth Israel Deaconness Medical Center

This study addresses the challenge of integrating randomized or single-arm clinical trials with external experimental or observational data to enable cross-study treatment comparisons and improve estimation precision of treatment effects. Methodologically, building upon the potential outcomes framework, we first develop a unified identification strategy for hybrid-data designs, systematically characterizing identifiability conditions across diverse designs—including historical controls, synthetic controls, and anchoring estimators—and propose a generalizable taxonomy of such designs along with corresponding causal inference principles. Our contribution lies in filling a critical theoretical gap in regulatory science regarding the rigorous integration of external controls, thereby establishing a methodological foundation for leveraging real-world evidence to complement trial-based evidence in pharmaceutical and medical device evaluation. This advancement significantly enhances the transportability of evidence and its applicability to regulatory decision-making.

Combining experimental studies with external data sourcesDeveloping identification strategies for treatment effectsFormalizing study designs to support systematic evaluation

Planning for Gold: Sample Splitting for Valid Powerful Design of Observational Studies

Jun 02, 2024
WB
William Bekerman
🏛️ University of Pennsylvania | World Bank

To address latent bias arising from unmeasured confounding in observational studies, this paper proposes a novel causal inference paradigm based on sample splitting: data are partitioned into planning and analysis samples, where the former adaptively selects robust design parameters (e.g., matching strategies, covariate sets), and the latter yields unbiased causal estimates. The method innovatively integrates multiple testing correction, heteroskedasticity-robust covariance estimation, and formal sensitivity analysis—extending support to multiple outcome variables, thereby relaxing the conventional single-outcome assumption. We establish theoretical guarantees of statistical validity under latent bias. Simulation studies demonstrate substantially higher statistical power than benchmark methods under strong unmeasured confounding. Empirical application to assessing the multidimensional impacts of floods in Bangladesh confirms practical feasibility and robustness.

Addresses bias from unmeasured covariates in observational studiesDevelops hypothesis screening method using split sample designEnables valid causal inference despite unmeasured confounding

Latest Papers

What's happening recently
View more

This study addresses the challenge of efficiently estimating causal effects under confounding when experimental budgets are limited. The authors propose a novel approach that integrates instrumental variable regression with Gaussian graphical models, leveraging prior knowledge of partial joint distributions to optimize the allocation between fully observed samples and partially observed data (e.g., only \(X_{12}\)). Under a fixed budget constraint, this method analytically derives the optimal sampling scheme that minimizes the asymptotic variance of the causal effect estimator—a solution not previously available in closed form. Theoretical analysis demonstrates that the proposed allocation significantly reduces both the total budget and the number of complete observations required to detect non-zero causal effects. Empirical validation in automotive analytics and drug discovery underscores the method’s practical utility alongside its theoretical contributions.

budget constraintcausal effect estimationexperimental design

This study addresses the limitations of traditional control-based causal inference methods—such as matching and difference-in-differences—in settings characterized by pervasive or structurally ambiguous spillover effects, where reliance on uncontaminated control units impedes accurate identification of both average direct and spillover effects. Within the potential outcomes framework, this work provides the first systematic comparison between control-based and prediction-based counterfactual approaches—including interrupted time series and machine learning control—in terms of their identification capabilities. Through simulation and empirical analyses, the authors demonstrate that in environments with widespread interference, prediction-based methods can more reliably estimate certain causal parameters over short horizons, circumventing the stringent assumption of unperturbed units and thereby offering a promising alternative for causal inference under complex interference.

causal inferencecounterfactualsinterference

This study addresses treatment switching in oncology randomized controlled trials, a phenomenon that violates randomization and introduces bias in overall survival estimation. To mitigate this issue, the authors propose a weighted causal inference framework that innovatively integrates external controls, synthetic controls, and balancing weights from observational studies. By incorporating multiple imputation and time-varying weights, the method effectively adjusts for the impact of treatment switching on efficacy estimates. The approach avoids strong parametric assumptions and complex modeling structures, demonstrating superior performance over conventional adjustment strategies in simulation studies. Its robustness and practical utility are further validated through applications to two phase III oncology clinical trials.

causal inferenceexternal controlsoverall survival

This study addresses the challenge of identifying heterogeneous treatment effects while controlling the false discovery rate (FDR) in matched observational studies, where limited sample sizes and unmeasured confounding pose significant obstacles. The authors propose a novel method that adaptively discovers interpretable subgroups defined by covariate thresholds under many-to-one matching designs. Their approach achieves exact FDR control at the subgroup level for the first time and integrates sensitivity analysis models to account for unobserved confounding, leveraging multiple controls to enhance statistical power. Theoretical analysis, simulations, and empirical evaluation demonstrate that the method outperforms existing baselines in both accuracy and power when estimating heterogeneous economic returns to college education.

effect modificationfalse discovery ratematched controls

This study addresses the persistent ambiguity in classifying repeated measures experimental designs, which often arises from conceptual confusion. To resolve this issue, the authors systematically clarify the core characteristics of such designs and propose a novel classification framework grounded in experimental units and randomization strategies. For the first time in this context, Hasse diagrams are introduced to visually represent the hierarchical structure of these designs. This approach effectively distinguishes among various types of repeated measures designs, eliminates terminological ambiguities, and substantially enhances both the rigor and interpretability of experimental planning and reporting.

experimental designexperimental unitsHasse diagrams

Hot Scholars

GN

Graham Neubig

Carnegie Mellon University, All Hands AI
Natural Language ProcessingMachine LearningArtificial Intelligence
IS

Ilia Sucholutsky

New York University
deep learningrepresentation learningsmall datarepresentational alignment
TL

Thomas L. Griffiths

Professor of Psychology and Computer Science, Princeton University
Computational Models of CognitionCognitive ScienceMachine LearningCognitive Psychology
PY

Pin-Yu Chen

Principal Research Scientist, IBM Research AI; MIT-IBM Watson AI Lab; RPI-IBM AIRC
AI SafetyGenerative AITrustworthy Machine LearningAdversarial Machine Learning
IT

Ivan Titov

University of Edinburgh / University of Amsterdam
natural language processingmachine learninginterpretability