Score
Designs and analyzes studies in which the same participants or items are measured under multiple conditions to estimate within-subject treatment or training effects; this includes constructing counterbalanced orders, matched-pair comparisons, randomizing pre/post sequences, and applying within-subject statistical analyses to control for between-subject variability and isolate causal differences.
Counterbalanced within-subject experiments risk invalid causal inference due to unverifiable and often violated assumptions—particularly the symmetry and cancelability of carryover effects. Method: We introduce “sequential exchangeability” as a formal identification assumption within the potential outcomes framework, rigorously exposing inherent limitations of counterbalancing; we then develop actionable strategies—including diagnostic tests, optimized washout periods, covariate adjustment, and alternative sequence designs—grounded in causal identification theory, sequential randomization modeling, and sensitivity analysis. Contribution/Results: Our work delineates precise validity boundaries for counterbalanced designs, providing rigorous, practical guidelines for within-subject experimentation in psychology, human-computer interaction, and related fields. By addressing foundational identifiability concerns, it substantially enhances the reliability of causal inference in repeated-measures settings.
This study addresses the challenge of integrating randomized or single-arm clinical trials with external experimental or observational data to enable cross-study treatment comparisons and improve estimation precision of treatment effects. Methodologically, building upon the potential outcomes framework, we first develop a unified identification strategy for hybrid-data designs, systematically characterizing identifiability conditions across diverse designs—including historical controls, synthetic controls, and anchoring estimators—and propose a generalizable taxonomy of such designs along with corresponding causal inference principles. Our contribution lies in filling a critical theoretical gap in regulatory science regarding the rigorous integration of external controls, thereby establishing a methodological foundation for leveraging real-world evidence to complement trial-based evidence in pharmaceutical and medical device evaluation. This advancement significantly enhances the transportability of evidence and its applicability to regulatory decision-making.
This paper addresses causal inference under violations of the Stable Unit Treatment Value Assumption (SUTVA) and interference among units. Method: We propose the Homogeneous-Intervention Average Treatment Effect (HAATE) as a new target estimand for the Global Average Treatment Effect (GATE); formally define HAATE; prove theoretically that the difference-in-means estimator dominates a correctly specified regression model under interference; and design a two-stage cluster-randomized experiment that leverages intra-cluster treatment correlation to model cluster-level error, thereby substantially reducing root mean squared error (RMSE). Contribution/Results: Monte Carlo simulations and a large-scale online A/B test on Facebook demonstrate that, compared to conventional designs, our approach significantly improves estimation accuracy in finite samples—enhancing the reliability of policy-level causal inference under interference.
Traditional randomized controlled trials (RCTs) suffer from cross-group spillovers and interference in multi-population interaction settings—e.g., buyer-seller or creator-subscriber systems—leading to biased estimation of the average treatment effect (ATE). This paper introduces the first systematic multidimensional randomization design framework, relaxing the conventional single-layer randomization assumption. By integrating hierarchical randomization, cross-group assignment, and potential outcomes modeling, it jointly identifies both the ATE and cross-group interference effects. We establish theoretical guarantees: the proposed unbiased estimator is consistent and asymptotically normal. Simulation results demonstrate substantially higher statistical power compared to standard RCTs. This work extends the scope of causal questions addressable through experimental design and provides a rigorous foundation for causal inference in complex intervention environments—particularly platform economies—where interdependent user behaviors induce non-negligible interference.
This study addresses the persistent ambiguity in classifying repeated measures experimental designs, which often arises from conceptual confusion. To resolve this issue, the authors systematically clarify the core characteristics of such designs and propose a novel classification framework grounded in experimental units and randomization strategies. For the first time in this context, Hasse diagrams are introduced to visually represent the hierarchical structure of these designs. This approach effectively distinguishes among various types of repeated measures designs, eliminates terminological ambiguities, and substantially enhances both the rigor and interpretability of experimental planning and reporting.
This study addresses the challenge of ensuring rigor in causal inference under multi-source heterogeneous data fusion by proposing a structured design paradigm grounded in the target trial framework. The approach explicitly incorporates the target population and its sampling model into the causal analysis, systematically integrating external controls, generalizability, and transportability assessments through data element alignment, transparent assumption articulation, and emulation of the target trial. Its key innovation lies in anchoring the entire framework to a precise definition of the target population, thereby identifying and mitigating irreconcilable conflicts across data sources. This strategy enhances both the reliability and interpretability of causal conclusions derived from complex, real-world data ecosystems.
This study addresses the challenge of accurately inferring the distribution of individual treatment effects—such as the proportion benefiting, the median effect, or the maximum impact—in randomized experiments, without suffering power loss due to suboptimal pre-specified test statistics. The authors propose an adaptive randomization test that combines multiple rank-based statistics, ensuring finite-sample validity without requiring prior knowledge of the optimal statistic. Innovatively integrating adaptive statistic combination with stratified weighting, the method effectively circumvents the power degradation typically induced by multiple comparison corrections and accommodates heterogeneous stratified experimental designs. In an empirical application to a teacher training program, the approach reveals that approximately half of the teachers experience significant benefits, demonstrating superior detection power and interpretability compared to conventional single rank-based tests.
This study addresses the challenge of identifying heterogeneous treatment effects while controlling the false discovery rate (FDR) in matched observational studies, where limited sample sizes and unmeasured confounding pose significant obstacles. The authors propose a novel method that adaptively discovers interpretable subgroups defined by covariate thresholds under many-to-one matching designs. Their approach achieves exact FDR control at the subgroup level for the first time and integrates sensitivity analysis models to account for unobserved confounding, leveraging multiple controls to enhance statistical power. Theoretical analysis, simulations, and empirical evaluation demonstrate that the method outperforms existing baselines in both accuracy and power when estimating heterogeneous economic returns to college education.