Score
Designing and analyzing experiments where the same subjects are measured under multiple conditions to control for between-subject variability, enabling comparisons such as cue effectiveness within a neglected visual field or changes in behavior when switching task formats. Covers assignment, counterbalancing, and statistical analysis appropriate for repeated-measures designs.
This study addresses the challenges of interpreting and analyzing high-dimensional experimental data within the traditional design of experiments (DoE) framework. By integrating analysis of variance (ANOVA) with simultaneous component analysis (SCA), the authors develop ASCA—a multivariate extension of ANOVA—and systematically unify a century of ANOVA and DoE theory to establish a rigorous application protocol tailored for high-dimensional data. Through a comprehensive literature review and illustrative case studies, the work defines a standardized analytical workflow and best practices that substantially enhance the interpretability and reliability of results from high-dimensional DoE studies. This contribution fills a critical methodological gap in the analysis of multivariate experimental data.
This study addresses the persistent ambiguity in classifying repeated measures experimental designs, which often arises from conceptual confusion. To resolve this issue, the authors systematically clarify the core characteristics of such designs and propose a novel classification framework grounded in experimental units and randomization strategies. For the first time in this context, Hasse diagrams are introduced to visually represent the hierarchical structure of these designs. This approach effectively distinguishes among various types of repeated measures designs, eliminates terminological ambiguities, and substantially enhances both the rigor and interpretability of experimental planning and reporting.
Counterbalanced within-subject experiments risk invalid causal inference due to unverifiable and often violated assumptions—particularly the symmetry and cancelability of carryover effects. Method: We introduce “sequential exchangeability” as a formal identification assumption within the potential outcomes framework, rigorously exposing inherent limitations of counterbalancing; we then develop actionable strategies—including diagnostic tests, optimized washout periods, covariate adjustment, and alternative sequence designs—grounded in causal identification theory, sequential randomization modeling, and sensitivity analysis. Contribution/Results: Our work delineates precise validity boundaries for counterbalanced designs, providing rigorous, practical guidelines for within-subject experimentation in psychology, human-computer interaction, and related fields. By addressing foundational identifiability concerns, it substantially enhances the reliability of causal inference in repeated-measures settings.
This study addresses the challenge of ensuring rigor in causal inference under multi-source heterogeneous data fusion by proposing a structured design paradigm grounded in the target trial framework. The approach explicitly incorporates the target population and its sampling model into the causal analysis, systematically integrating external controls, generalizability, and transportability assessments through data element alignment, transparent assumption articulation, and emulation of the target trial. Its key innovation lies in anchoring the entire framework to a precise definition of the target population, thereby identifying and mitigating irreconcilable conflicts across data sources. This strategy enhances both the reliability and interpretability of causal conclusions derived from complex, real-world data ecosystems.
Traditional randomized controlled trials (RCTs) suffer from cross-group spillovers and interference in multi-population interaction settings—e.g., buyer-seller or creator-subscriber systems—leading to biased estimation of the average treatment effect (ATE). This paper introduces the first systematic multidimensional randomization design framework, relaxing the conventional single-layer randomization assumption. By integrating hierarchical randomization, cross-group assignment, and potential outcomes modeling, it jointly identifies both the ATE and cross-group interference effects. We establish theoretical guarantees: the proposed unbiased estimator is consistent and asymptotically normal. Simulation results demonstrate substantially higher statistical power compared to standard RCTs. This work extends the scope of causal questions addressable through experimental design and provides a rigorous foundation for causal inference in complex intervention environments—particularly platform economies—where interdependent user behaviors induce non-negligible interference.
The reproducibility crisis in psychology is partly attributable to uncorrected multiple comparisons, inflating false discovery rates and contributing to replication failures. This study provides the first systematic quantification of how multiple comparisons contributed to non-replication across 88 psychological studies and introduces TreeBH—a novel false discovery rate (FDR) control procedure designed for hierarchical hypothesis structures. Empirical evaluation shows that applying TreeBH renders 21 originally significant findings non-significant; of these, 20 were indeed not replicated in follow-up studies—accounting for 34% of all non-replicated results—while preserving 97% statistical power. This work establishes uncorrected multiple comparisons as a key driver of the reproducibility crisis and delivers the first theoretically rigorous, empirically feasible hierarchical multiple testing correction framework tailored to typical experimental designs in psychology.
This study addresses the challenge of biased effect estimation in online controlled experiments caused by overlapping tests on shared traffic, which hinders accurate assessment of feature interactions. To resolve this, the authors propose Multi-Experiment Analysis (MEA), a method grounded in statistical modeling and causal inference that consistently estimates joint effects under arbitrary partial or full overlap and multi-variant settings—without requiring predefined factorial designs or constrained traffic allocation. MEA uniquely enables, without coordination overhead, the simultaneous modeling of bias-corrected individual effects, joint effects for any combination of variants, and conditional effects. Simulations confirm the estimator’s consistency and nominal confidence interval coverage, and the approach has been successfully deployed in large-scale production systems across multiple real-world business applications.
This study addresses the lack of systematic support for creating, managing, and deploying stimulus materials in visualization experiments—a gap that often leads to invalid results or wasted resources. Through semi-structured interviews with 19 visualization researchers, the work systematically examines practices and challenges across the full lifecycle of stimulus materials, from exploration and selection to deployment and analysis, integrating perspectives from user research and human factors engineering. The findings reveal, for the first time, a heavy reliance on manual processes and significant scalability limitations as core pain points in current workflows. Building on these insights, the study identifies key opportunities for improvement, including automated generation and intelligent validation of stimuli, thereby laying the groundwork for future directions such as AI-assisted stimulus design.
This study addresses the underappreciated challenge of estimation and communication following multiplicity adjustment within the frequentist framework in complex clinical trials, where multiple endpoints, interim data looks, or group comparisons often introduce estimation bias and complicate interpretation, thereby undermining transparency in benefit–risk assessment. By integrating advanced methodologies such as adaptive designs and graphical approaches to multiple testing, the work illustrates through concrete examples the limitations of current strategies in conveying trial results meaningfully. The research underscores the need to critically reevaluate prevailing practices and foster interdisciplinary dialogue to enhance both the accuracy of effect estimation and the clarity of result communication, ultimately informing future methodological standards and regulatory guidance.
This work addresses the lack of a systematic framework to guide experimental design decisions in replication studies. It proposes the first multidimensional design space framework specifically tailored for replication research, conceptualizing replication as a pairwise comparison problem. The framework structures replication planning and analysis through four practical dimensions—task, data, method, and metrics—and three comparative levels: micro, meso, and macro. By offering actionable design guidelines and a comprehensive taxonomy, it enables both prospective planning and retrospective evaluation of replication efforts. Empirical case studies in visualization and human-computer interaction demonstrate the framework’s effectiveness in enhancing the rigor of replication designs and improving the comparability of evaluation outcomes.