Score
Designs and conducts controlled comparative studies and experiments, and builds statistical and econometric comparison models (e.g., difference‑in‑differences and pairwise comparison models) to measure and test differences between systems, treatments, models, or conditions. Analyzes comparative outcomes, equilibria, risks, and robustness using comparative statistical and experimental methods and synthesizes findings across studies and literature to evaluate relative performance and draw causal or descriptive inferences.
In methodological comparative studies, algorithmic failures—such as non-convergence or absence of output—preclude performance evaluation, yet existing literature lacks standardized guidelines for handling such failures, often overlooking or misapplying failure mitigation strategies. Method: We systematically analyze failure causes and risks of improper handling, critically examine prevalent censoring and imputation strategies for their statistical biases, and propose the principle of “context-adapted failure fallback,” establishing a framework grounded in empirically feasible fallback mechanisms. Through statistical modeling, failure root-cause diagnosis, and cross-domain empirical analysis, we identify widespread deficiencies in published studies’ failure handling practices. Contribution/Results: Two representative case studies demonstrate that inappropriate failure handling significantly distorts method rankings and undermines conclusion validity. Our work bridges critical theoretical and practical gaps in the principled treatment of algorithmic failures in empirical methodology research.
Estimating causal effects from observational data requires selecting appropriate control and instrumental variables that satisfy causal identification conditions—a challenging task often reliant on strong domain knowledge or ad hoc assumptions. Method: This paper proposes the first end-to-end joint learning framework that automatically identifies valid combinations of control and instrumental variables. Grounded in conditional independence testing, the method integrates nonparametric dependence measures with structural search optimization, ensuring statistical consistency in variable selection under mild regularity conditions. Contribution/Results: Unlike conventional approaches requiring prespecified variable sets or strong prior assumptions, our framework is fully data-driven. In simulations, it achieves significantly higher variable identification accuracy. Empirically, applied to the Job Corps study, its estimated treatment effect closely aligns with results from the randomized controlled trial—demonstrating both validity and robustness in real-world causal inference.
In large-scale aggregate-unit experiments (e.g., markets), conventional randomized treatment assignment often yields severe baseline imbalance due to extremely few treated units, leading to biased causal estimates. To address this, we systematically integrate the synthetic control method into experimental design, proposing a non-randomized treatment allocation mechanism: dynamically constructing a weighted synthetic control group based on pre-treatment covariates. We further develop配套 components—including counterfactual prediction, distance-driven unit matching, robust variance estimation, and a novel confidence interval construction procedure. Theoretically, our estimator is proven consistent and asymptotically normal. Empirically, it reduces estimation bias by 40–65% relative to standard randomization and substantially improves statistical power. Our core contribution is a new causal inference paradigm for small-N aggregate experiments—rigorous in inference, unbiased under mild assumptions, and highly interpretable.
This study addresses the challenge of integrating randomized or single-arm clinical trials with external experimental or observational data to enable cross-study treatment comparisons and improve estimation precision of treatment effects. Methodologically, building upon the potential outcomes framework, we first develop a unified identification strategy for hybrid-data designs, systematically characterizing identifiability conditions across diverse designs—including historical controls, synthetic controls, and anchoring estimators—and propose a generalizable taxonomy of such designs along with corresponding causal inference principles. Our contribution lies in filling a critical theoretical gap in regulatory science regarding the rigorous integration of external controls, thereby establishing a methodological foundation for leveraging real-world evidence to complement trial-based evidence in pharmaceutical and medical device evaluation. This advancement significantly enhances the transportability of evidence and its applicability to regulatory decision-making.
Existing R tools lack systematic support for crossover design data—particularly those involving longitudinal within-period measurements and carryover effects. This paper introduces CrossCarry, the first open-source R package specifically designed for modeling crossover trials. It accommodates exponential-family responses, arbitrary-order designs, and scenarios with or without washout periods. Methodologically: (1) it extends the generalized estimating equations (GEE) framework by jointly modeling within-period correlation structures and between-period carryover dependencies—a novel integration; (2) it incorporates B-spline–based nonparametric components to flexibly estimate both temporal trends and carryover effects; and (3) it enables unified, flexible modeling of treatment, time, and carryover effects. Empirical evaluations demonstrate that CrossCarry substantially improves statistical power and estimation accuracy for treatment and carryover effects under challenging conditions—including skewed responses and weak washout.
Estimating the functional relationship between a continuous exposure and a binary outcome is challenging when covariates are measured with error. This study presents the first systematic evaluation of Simulation-Extrapolation, Regression Calibration, multiple imputation, and Bayesian correction methods, each coupled with flexible modeling techniques—including B-splines, P-splines, and fractional polynomials—within a multi-team, fully blinded, neutral simulation framework. By generating 155 distinct simulation scenarios and repeated samples, the research quantifies the bias and variance of each approach, revealing their relative strengths and limitations. The findings not only inform method selection under measurement error but also demonstrate the feasibility and value of this neutral comparative paradigm for rigorous methodological assessment.
This study challenges the conventional wisdom that representative sampling is always optimal in randomized controlled trials (RCTs) with limited budgets. It proposes a novel optimal sampling framework that integrates cost structures and prior knowledge of heterogeneous treatment effects. By combining Bayesian priors, heterogeneity modeling, hypothesis testing, and expected utility optimization, the authors theoretically demonstrate that under tight budget constraints, concentrating sampling efforts on a single high-potential subpopulation can substantially enhance the expected impact of downstream interventions. Only when the budget is sufficiently large does the optimal strategy converge to representative sampling. These findings hold across diverse resource-constrained experimental settings and offer a new paradigm for efficient, evidence-based decision-making in applied research.
This work addresses the inefficiencies in clinical trial design and analysis that hinder drug development success rates and timelines. It proposes establishing a dedicated statistical methodology team within pharmaceutical companies as a strategic investment, embedded through a systematic organizational structure and cross-functional collaboration mechanisms—both internally across departments and externally with academic and regulatory partners—to break down information silos. By integrating advanced statistical modeling, optimized clinical trial designs, and other high-impact quantitative methodologies, this team significantly enhances R&D efficiency, shortens development cycles, and strengthens the scientific rigor and likelihood of success in clinical decision-making.
Traditional hybrid experimental designs struggle to robustly control the frequentist operating characteristics of Bayesian decisions under model misspecification and lack efficient sample size determination methods applicable to generalized posteriors. This work proposes a computationally efficient experimental design framework that requires simulations at only two sample sizes and leverages extrapolation modeling of posterior summary functions to infer performance across the entire sample size space. This approach enables identification of the minimal sample size and decision rule satisfying desired operating characteristics. It represents the first general and scalable method for sample size planning under generalized posteriors, substantially reducing computational burden while enhancing robustness to model misspecification. The method’s validity and broad applicability within Bayesian M-estimation–type experiments are demonstrated through the redesign of an adaptive clinical trial with time-to-event outcomes.