within-subject comparison

Designing and analyzing experiments where the same subjects are measured under multiple conditions to control for between-subject variability, enabling comparisons such as cue effectiveness within a neglected visual field or changes in behavior when switching task formats. Covers assignment, counterbalancing, and statistical analysis appropriate for repeated-measures designs.

within-subjectcomparison

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the persistent ambiguity in classifying repeated measures experimental designs, which often arises from conceptual confusion. To resolve this issue, the authors systematically clarify the core characteristics of such designs and propose a novel classification framework grounded in experimental units and randomization strategies. For the first time in this context, Hasse diagrams are introduced to visually represent the hierarchical structure of these designs. This approach effectively distinguishes among various types of repeated measures designs, eliminates terminological ambiguities, and substantially enhances both the rigor and interpretability of experimental planning and reporting.

experimental designexperimental unitsHasse diagrams

Causal Inference in Counterbalanced Within-Subjects Designs

May 06, 2025
JH
Justin Ho
🏛️ Harvard University | University of California, Berkeley

Counterbalanced within-subject experiments risk invalid causal inference due to unverifiable and often violated assumptions—particularly the symmetry and cancelability of carryover effects. Method: We introduce “sequential exchangeability” as a formal identification assumption within the potential outcomes framework, rigorously exposing inherent limitations of counterbalancing; we then develop actionable strategies—including diagnostic tests, optimized washout periods, covariate adjustment, and alternative sequence designs—grounded in causal identification theory, sequential randomization modeling, and sensitivity analysis. Contribution/Results: Our work delineates precise validity boundaries for counterbalanced designs, providing rigorous, practical guidelines for within-subject experimentation in psychology, human-computer interaction, and related fields. By addressing foundational identifiability concerns, it substantially enhances the reliability of causal inference in repeated-measures settings.

Challenges in causal inference with counterbalanced within-subjects designsNeed for alternative methods to ensure valid causal inferenceUnverifiable assumptions about symmetric carryover effects in counterbalancing

This study addresses the challenge of ensuring rigor in causal inference under multi-source heterogeneous data fusion by proposing a structured design paradigm grounded in the target trial framework. The approach explicitly incorporates the target population and its sampling model into the causal analysis, systematically integrating external controls, generalizability, and transportability assessments through data element alignment, transparent assumption articulation, and emulation of the target trial. Its key innovation lies in anchoring the entire framework to a precise definition of the target population, thereby identifying and mitigating irreconcilable conflicts across data sources. This strategy enhances both the reliability and interpretability of causal conclusions derived from complex, real-world data ecosystems.

causal inferencedata integrationexternal comparator analyses

Multiple Randomization Designs

Dec 27, 2021
LM
Lorenzo Masoero
🏛️ Amazon | University of Washington | Stanford University

Traditional randomized controlled trials (RCTs) suffer from cross-group spillovers and interference in multi-population interaction settings—e.g., buyer-seller or creator-subscriber systems—leading to biased estimation of the average treatment effect (ATE). This paper introduces the first systematic multidimensional randomization design framework, relaxing the conventional single-layer randomization assumption. By integrating hierarchical randomization, cross-group assignment, and potential outcomes modeling, it jointly identifies both the ATE and cross-group interference effects. We establish theoretical guarantees: the proposed unbiased estimator is consistent and asymptotically normal. Simulation results demonstrate substantially higher statistical power compared to standard RCTs. This work extends the scope of causal questions addressable through experimental design and provides a rigorous foundation for causal inference in complex intervention environments—particularly platform economies—where interdependent user behaviors induce non-negligible interference.

Addressing interference effects in multi-population experimental settingsDeveloping statistical methods for analyzing multiple randomization designsProposing new designs for experiments with interacting populations

Time to adjust: Improving replicability in experimental psychology by adjustment for evident selective inference

Jun 20, 2020
YZ
Yoav Zeevi
🏛️ Tel Aviv University | Canadian Institute for Advanced Research

The reproducibility crisis in psychology is partly attributable to uncorrected multiple comparisons, inflating false discovery rates and contributing to replication failures. This study provides the first systematic quantification of how multiple comparisons contributed to non-replication across 88 psychological studies and introduces TreeBH—a novel false discovery rate (FDR) control procedure designed for hierarchical hypothesis structures. Empirical evaluation shows that applying TreeBH renders 21 originally significant findings non-significant; of these, 20 were indeed not replicated in follow-up studies—accounting for 34% of all non-replicated results—while preserving 97% statistical power. This work establishes uncorrected multiple comparisons as a key driver of the reproducibility crisis and delivers the first theoretically rigorous, empirically feasible hierarchical multiple testing correction framework tailored to typical experimental designs in psychology.

Addressing replicability crisis in psychology by adjusting for selective inferenceHighlighting overlooked issue of unadjusted multiple comparisons in studiesProposing hierarchical FDR adjustment to improve experimental replicability

Latest Papers

What's happening recently
View more

This study addresses the challenge of biased effect estimation in online controlled experiments caused by overlapping tests on shared traffic, which hinders accurate assessment of feature interactions. To resolve this, the authors propose Multi-Experiment Analysis (MEA), a method grounded in statistical modeling and causal inference that consistently estimates joint effects under arbitrary partial or full overlap and multi-variant settings—without requiring predefined factorial designs or constrained traffic allocation. MEA uniquely enables, without coordination overhead, the simultaneous modeling of bias-corrected individual effects, joint effects for any combination of variants, and conditional effects. Simulations confirm the estimator’s consistency and nominal confidence interval coverage, and the approach has been successfully deployed in large-scale production systems across multiple real-world business applications.

experiment overlapfeature interactiononline controlled experiments

This study addresses the lack of systematic support for creating, managing, and deploying stimulus materials in visualization experiments—a gap that often leads to invalid results or wasted resources. Through semi-structured interviews with 19 visualization researchers, the work systematically examines practices and challenges across the full lifecycle of stimulus materials, from exploration and selection to deployment and analysis, integrating perspectives from user research and human factors engineering. The findings reveal, for the first time, a heavy reliance on manual processes and significant scalability limitations as core pain points in current workflows. Building on these insights, the study identifies key opportunities for improvement, including automated generation and intelligent validation of stimuli, thereby laying the groundwork for future directions such as AI-assisted stimulus design.

experimental designresearch challengesstimuli deployment

This study addresses the underappreciated challenge of estimation and communication following multiplicity adjustment within the frequentist framework in complex clinical trials, where multiple endpoints, interim data looks, or group comparisons often introduce estimation bias and complicate interpretation, thereby undermining transparency in benefit–risk assessment. By integrating advanced methodologies such as adaptive designs and graphical approaches to multiple testing, the work illustrates through concrete examples the limitations of current strategies in conveying trial results meaningfully. The research underscores the need to critically reevaluate prevailing practices and foster interdisciplinary dialogue to enhance both the accuracy of effect estimation and the clarity of result communication, ultimately informing future methodological standards and regulatory guidance.

adaptive designestimationmultiple hypotheses

This work addresses the lack of a systematic framework to guide experimental design decisions in replication studies. It proposes the first multidimensional design space framework specifically tailored for replication research, conceptualizing replication as a pairwise comparison problem. The framework structures replication planning and analysis through four practical dimensions—task, data, method, and metrics—and three comparative levels: micro, meso, and macro. By offering actionable design guidelines and a comprehensive taxonomy, it enables both prospective planning and retrospective evaluation of replication efforts. Empirical case studies in visualization and human-computer interaction demonstrate the framework’s effectiveness in enhancing the rigor of replication designs and improving the comparability of evaluation outcomes.

design spaceexperimental designHCI

Hot Scholars

AO

Alice Oh

KAIST Computer Science
machine learningNLPcomputational social science
YP

Yujin Park

Georgia Southern University
Online LearningTeacher Professional LearningElementary STEMDigital Literacy
HC

Haejun Chung

Associate Professor, Hanyang University
ElectromagneticsMetasurfaceOpticsPhotonics
IJ

Ikbeom Jang

MGH/Harvard Medical School
Medical ImagingMachine LearningBrainImaging Biomarker
MS

Markus Strohmaier

University of Mannheim, GESIS - Leibniz Institute for the Social Sciences, CSH Vienna
Computational Social ScienceComputational Social SystemsSocial Data Science