within-subject evaluation

Designs and analyzes studies in which the same participants or items are measured under multiple conditions to estimate within-subject treatment or training effects; this includes constructing counterbalanced orders, matched-pair comparisons, randomizing pre/post sequences, and applying within-subject statistical analyses to control for between-subject variability and isolate causal differences.

within-subjectevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$207K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Causal Inference in Counterbalanced Within-Subjects Designs

May 06, 2025
JH
Justin Ho
🏛️ Harvard University | University of California, Berkeley

Counterbalanced within-subject experiments risk invalid causal inference due to unverifiable and often violated assumptions—particularly the symmetry and cancelability of carryover effects. Method: We introduce “sequential exchangeability” as a formal identification assumption within the potential outcomes framework, rigorously exposing inherent limitations of counterbalancing; we then develop actionable strategies—including diagnostic tests, optimized washout periods, covariate adjustment, and alternative sequence designs—grounded in causal identification theory, sequential randomization modeling, and sensitivity analysis. Contribution/Results: Our work delineates precise validity boundaries for counterbalanced designs, providing rigorous, practical guidelines for within-subject experimentation in psychology, human-computer interaction, and related fields. By addressing foundational identifiability concerns, it substantially enhances the reliability of causal inference in repeated-measures settings.

Challenges in causal inference with counterbalanced within-subjects designsNeed for alternative methods to ensure valid causal inferenceUnverifiable assumptions about symmetric carryover effects in counterbalancing

Combining an experimental study with external data: study designs and identification strategies

Jun 05, 2024
LU
Lawson Ung
🏛️ Harvard T.H. Chan School of Public Health | Dartmouth Geisel School of Medicine | Beth Israel Deaconness Medical Center

This study addresses the challenge of integrating randomized or single-arm clinical trials with external experimental or observational data to enable cross-study treatment comparisons and improve estimation precision of treatment effects. Methodologically, building upon the potential outcomes framework, we first develop a unified identification strategy for hybrid-data designs, systematically characterizing identifiability conditions across diverse designs—including historical controls, synthetic controls, and anchoring estimators—and propose a generalizable taxonomy of such designs along with corresponding causal inference principles. Our contribution lies in filling a critical theoretical gap in regulatory science regarding the rigorous integration of external controls, thereby establishing a methodological foundation for leveraging real-world evidence to complement trial-based evidence in pharmaceutical and medical device evaluation. This advancement significantly enhances the transportability of evidence and its applicability to regulatory decision-making.

Combining experimental studies with external data sourcesDeveloping identification strategies for treatment effectsFormalizing study designs to support systematic evaluation

Experimentation for Homogenous Policy Change

Jan 28, 2021
MO
Molly Offer-Westort

This paper addresses causal inference under violations of the Stable Unit Treatment Value Assumption (SUTVA) and interference among units. Method: We propose the Homogeneous-Intervention Average Treatment Effect (HAATE) as a new target estimand for the Global Average Treatment Effect (GATE); formally define HAATE; prove theoretically that the difference-in-means estimator dominates a correctly specified regression model under interference; and design a two-stage cluster-randomized experiment that leverages intra-cluster treatment correlation to model cluster-level error, thereby substantially reducing root mean squared error (RMSE). Contribution/Results: Monte Carlo simulations and a large-scale online A/B test on Facebook demonstrate that, compared to conventional designs, our approach significantly improves estimation accuracy in finite samples—enhancing the reliability of policy-level causal inference under interference.

Comparing performance of estimators for Global Average Treatment EffectsEstimating treatment effects under interference among unitsEvaluating randomization designs for cluster-level correlated errors

Multiple Randomization Designs

Dec 27, 2021
LM
Lorenzo Masoero
🏛️ Amazon | University of Washington | Stanford University

Traditional randomized controlled trials (RCTs) suffer from cross-group spillovers and interference in multi-population interaction settings—e.g., buyer-seller or creator-subscriber systems—leading to biased estimation of the average treatment effect (ATE). This paper introduces the first systematic multidimensional randomization design framework, relaxing the conventional single-layer randomization assumption. By integrating hierarchical randomization, cross-group assignment, and potential outcomes modeling, it jointly identifies both the ATE and cross-group interference effects. We establish theoretical guarantees: the proposed unbiased estimator is consistent and asymptotically normal. Simulation results demonstrate substantially higher statistical power compared to standard RCTs. This work extends the scope of causal questions addressable through experimental design and provides a rigorous foundation for causal inference in complex intervention environments—particularly platform economies—where interdependent user behaviors induce non-negligible interference.

Addressing interference effects in multi-population experimental settingsDeveloping statistical methods for analyzing multiple randomization designsProposing new designs for experiments with interacting populations

Latest Papers

What's happening recently
View more

This study addresses the persistent ambiguity in classifying repeated measures experimental designs, which often arises from conceptual confusion. To resolve this issue, the authors systematically clarify the core characteristics of such designs and propose a novel classification framework grounded in experimental units and randomization strategies. For the first time in this context, Hasse diagrams are introduced to visually represent the hierarchical structure of these designs. This approach effectively distinguishes among various types of repeated measures designs, eliminates terminological ambiguities, and substantially enhances both the rigor and interpretability of experimental planning and reporting.

experimental designexperimental unitsHasse diagrams

This study addresses the challenge of ensuring rigor in causal inference under multi-source heterogeneous data fusion by proposing a structured design paradigm grounded in the target trial framework. The approach explicitly incorporates the target population and its sampling model into the causal analysis, systematically integrating external controls, generalizability, and transportability assessments through data element alignment, transparent assumption articulation, and emulation of the target trial. Its key innovation lies in anchoring the entire framework to a precise definition of the target population, thereby identifying and mitigating irreconcilable conflicts across data sources. This strategy enhances both the reliability and interpretability of causal conclusions derived from complex, real-world data ecosystems.

causal inferencedata integrationexternal comparator analyses

This study addresses the challenge of accurately inferring the distribution of individual treatment effects—such as the proportion benefiting, the median effect, or the maximum impact—in randomized experiments, without suffering power loss due to suboptimal pre-specified test statistics. The authors propose an adaptive randomization test that combines multiple rank-based statistics, ensuring finite-sample validity without requiring prior knowledge of the optimal statistic. Innovatively integrating adaptive statistic combination with stratified weighting, the method effectively circumvents the power degradation typically induced by multiple comparison corrections and accommodates heterogeneous stratified experimental designs. In an empirical application to a teacher training program, the approach reveals that approximately half of the teachers experience significant benefits, demonstrating superior detection power and interpretability compared to conventional single rank-based tests.

distributional inferenceindividual treatment effectsrandomization tests

This study addresses the challenge of identifying heterogeneous treatment effects while controlling the false discovery rate (FDR) in matched observational studies, where limited sample sizes and unmeasured confounding pose significant obstacles. The authors propose a novel method that adaptively discovers interpretable subgroups defined by covariate thresholds under many-to-one matching designs. Their approach achieves exact FDR control at the subgroup level for the first time and integrates sensitivity analysis models to account for unobserved confounding, leveraging multiple controls to enhance statistical power. Theoretical analysis, simulations, and empirical evaluation demonstrate that the method outperforms existing baselines in both accuracy and power when estimating heterogeneous economic returns to college education.

effect modificationfalse discovery ratematched controls

Hot Scholars

MV

Maurizio Vergari

Research Assistant, TU Berlin
UX for Extended Reality (XR)
XM

Xiaojuan Ma

Hong Kong University of Science and Technology
Human-Computer InteractionHuman-Engaged ComputingAffective Computing
AE

Abdallah El Ali

Centrum Wiskunde & Informatica, Utrecht University
Human Computer InteractionAffective ComputingMultimodal InteractionVR/AR/MR
MC

Mark Colley

University College London
Automated DrivingAugmented RealityDriver-Vehicle InteractionAccessibility