Score
Designs, builds, and analyzes experimental protocols and procedures — including randomized and controlled experiments, observational and field studies, intervention and reader/user studies, and probing and RL-specific experiments — that define study populations or samples, control or comparison conditions, interventions and manipulations, temporal stimulus sequences, and outcome measures or endpoints. Plans data collection and management, computes sample size and statistical power, implements random assignment and experimental controls, and evaluates experimental analysis outputs such as causal effect estimates, false positive rates, standard errors, and model-discrimination/sample-efficiency metrics.
This study addresses the challenge of integrating randomized or single-arm clinical trials with external experimental or observational data to enable cross-study treatment comparisons and improve estimation precision of treatment effects. Methodologically, building upon the potential outcomes framework, we first develop a unified identification strategy for hybrid-data designs, systematically characterizing identifiability conditions across diverse designs—including historical controls, synthetic controls, and anchoring estimators—and propose a generalizable taxonomy of such designs along with corresponding causal inference principles. Our contribution lies in filling a critical theoretical gap in regulatory science regarding the rigorous integration of external controls, thereby establishing a methodological foundation for leveraging real-world evidence to complement trial-based evidence in pharmaceutical and medical device evaluation. This advancement significantly enhances the transportability of evidence and its applicability to regulatory decision-making.
In large-scale aggregate-unit experiments (e.g., markets), conventional randomized treatment assignment often yields severe baseline imbalance due to extremely few treated units, leading to biased causal estimates. To address this, we systematically integrate the synthetic control method into experimental design, proposing a non-randomized treatment allocation mechanism: dynamically constructing a weighted synthetic control group based on pre-treatment covariates. We further develop配套 components—including counterfactual prediction, distance-driven unit matching, robust variance estimation, and a novel confidence interval construction procedure. Theoretically, our estimator is proven consistent and asymptotically normal. Empirically, it reduces estimation bias by 40–65% relative to standard randomization and substantially improves statistical power. Our core contribution is a new causal inference paradigm for small-N aggregate experiments—rigorous in inference, unbiased under mild assumptions, and highly interpretable.
Traditional randomized controlled trials (RCTs) suffer from cross-group spillovers and interference in multi-population interaction settings—e.g., buyer-seller or creator-subscriber systems—leading to biased estimation of the average treatment effect (ATE). This paper introduces the first systematic multidimensional randomization design framework, relaxing the conventional single-layer randomization assumption. By integrating hierarchical randomization, cross-group assignment, and potential outcomes modeling, it jointly identifies both the ATE and cross-group interference effects. We establish theoretical guarantees: the proposed unbiased estimator is consistent and asymptotically normal. Simulation results demonstrate substantially higher statistical power compared to standard RCTs. This work extends the scope of causal questions addressable through experimental design and provides a rigorous foundation for causal inference in complex intervention environments—particularly platform economies—where interdependent user behaviors induce non-negligible interference.
To address latent bias arising from unmeasured confounding in observational studies, this paper proposes a novel causal inference paradigm based on sample splitting: data are partitioned into planning and analysis samples, where the former adaptively selects robust design parameters (e.g., matching strategies, covariate sets), and the latter yields unbiased causal estimates. The method innovatively integrates multiple testing correction, heteroskedasticity-robust covariance estimation, and formal sensitivity analysis—extending support to multiple outcome variables, thereby relaxing the conventional single-outcome assumption. We establish theoretical guarantees of statistical validity under latent bias. Simulation studies demonstrate substantially higher statistical power than benchmark methods under strong unmeasured confounding. Empirical application to assessing the multidimensional impacts of floods in Bangladesh confirms practical feasibility and robustness.
This paper identifies a causal inference bias in service-intervention randomized controlled trials (RCTs) arising from provider capacity constraints: limited resources induce cross-participant interference and treatment dose heterogeneity, rendering the average treatment effect dependent on both sample size and capacity, and causing effect attenuation beyond a critical threshold—yielding non-monotonic, inverted-U-shaped statistical power. We formally characterize this capacity-constrained, queue-based interference mechanism for the first time, integrating queuing theory (square-root staffing rule), causal inference, and experimental design. Our method jointly optimizes provider capacity and sample size to maximize power under resource constraints. Results demonstrate substantial gains in statistical power, reduced required resources and participant enrollment, and provide a mechanistic explanation for the common phenomenon of intervention efficacy fading upon scaling—from RCT success to real-world failure.
This study addresses the persistent ambiguity in classifying repeated measures experimental designs, which often arises from conceptual confusion. To resolve this issue, the authors systematically clarify the core characteristics of such designs and propose a novel classification framework grounded in experimental units and randomization strategies. For the first time in this context, Hasse diagrams are introduced to visually represent the hierarchical structure of these designs. This approach effectively distinguishes among various types of repeated measures designs, eliminates terminological ambiguities, and substantially enhances both the rigor and interpretability of experimental planning and reporting.
Traditional hybrid experimental designs struggle to robustly control the frequentist operating characteristics of Bayesian decisions under model misspecification and lack efficient sample size determination methods applicable to generalized posteriors. This work proposes a computationally efficient experimental design framework that requires simulations at only two sample sizes and leverages extrapolation modeling of posterior summary functions to infer performance across the entire sample size space. This approach enables identification of the minimal sample size and decision rule satisfying desired operating characteristics. It represents the first general and scalable method for sample size planning under generalized posteriors, substantially reducing computational burden while enhancing robustness to model misspecification. The method’s validity and broad applicability within Bayesian M-estimation–type experiments are demonstrated through the redesign of an adaptive clinical trial with time-to-event outcomes.
This study challenges the conventional wisdom that representative sampling is always optimal in randomized controlled trials (RCTs) with limited budgets. It proposes a novel optimal sampling framework that integrates cost structures and prior knowledge of heterogeneous treatment effects. By combining Bayesian priors, heterogeneity modeling, hypothesis testing, and expected utility optimization, the authors theoretically demonstrate that under tight budget constraints, concentrating sampling efforts on a single high-potential subpopulation can substantially enhance the expected impact of downstream interventions. Only when the budget is sufficiently large does the optimal strategy converge to representative sampling. These findings hold across diverse resource-constrained experimental settings and offer a new paradigm for efficient, evidence-based decision-making in applied research.
This study addresses the challenge that data-driven protocol selection in target trial emulation often invalidates statistical inference. To resolve this, the authors propose a two-stage strategy based on sample splitting: the first subsample is used to explore and finalize the target trial protocol, while the second, independent subsample is employed to implement the selected protocol and conduct causal inference. Inspired by the exploratory-to-confirmatory paradigm in clinical trials, this approach decouples protocol specification from inference, thereby preserving the flexibility of scientific exploration while rigorously maintaining nominal coverage guarantees for statistical inference. By integrating sample splitting, target trial emulation, and causal inference, the work provides both theoretical assurance and a practical framework for valid statistical inference in observational studies.