Score
Designs and conducts analyses that determine whether and how causal effect estimates obtained in one setting (source) can be validly carried over to another setting (target), by specifying selection mechanisms, deriving transportability conditions and transport formulas, and identifying required adjustment sets. Builds estimators and sensitivity analyses or assessments that implement those conditions, and analyzes which assumptions or measured variables must hold or be collected to enable or evaluate transport of causal conclusions across environments.
This study addresses the practical challenge in external validity assessment where only aggregate data (AgD) from the target population are available and outcomes are subject to right-censoring. We propose TADA—the first two-stage weighted framework specifically designed for AgD—that jointly corrects for censoring bias and imbalance in effect modifier distributions without requiring individual patient data (IPD) from the source population. TADA integrates inverse probability censoring weighting (IPCW) with a moment-based participation weighting method. Simulation results demonstrate that, under typical censoring rates, TADA reduces extrapolation bias in treatment effect estimation by over 60% compared to conventional approaches. Consequently, TADA substantially enhances the feasibility, robustness, and clinical interpretability of transportability analyses in real-world settings where IPD are inaccessible or unavailable.
This study addresses the joint identifiability and transportability of causal effects across populations under unknown causal graphs. Motivated by the clinical challenge that experimental data from a source population often cannot be directly generalized to a target population, we propose the first Bayesian framework that does not require prior knowledge of the causal graph—unifying identifiability and transportability assessment within a single probabilistic model. Our method performs joint likelihood estimation using randomized trial data from the source domain and observational data from the target domain, grounded in mild assumptions from structural causal models. It enables probabilistic determination and joint inference of Z-conditional causal effects. Simulation results demonstrate that our approach accurately identifies transportable effects and significantly improves both unbiasedness and statistical efficiency in estimating target-population causal effects compared to single-source estimators—thereby overcoming the traditional reliance on known causal graph structures.
Causal effect estimation suffers from uncertainty in covariate selection—particularly under non-unique graph representations such as CPDAGs—and degraded accuracy in small-sample settings. Method: This paper proposes a practical, theoretically grounded criterion and algorithm for optimal adjustment set selection applicable to both DAGs and CPDAGs. It establishes, for the first time, a computability theorem for causal effects under CPDAGs, and designs an adjustment set selection criterion that jointly ensures identifiability and statistical efficiency, integrating the backdoor criterion, graph structure search, and rigorous theoretical justification. Results: Experiments on synthetic and real-world datasets demonstrate substantial improvements in causal effect estimation accuracy. The method exhibits strong robustness under small-sample conditions and when confounding structures are uncertain, providing a reliable, interpretable framework for covariate selection in data-limited causal inference.
This study addresses the limited external validity of software engineering experiments, often stemming from unrepresentative samples. It pioneers the systematic application of causal inference–based transportability methods in this domain, integrating experimental and observational data to develop tailored implementation pathways and practical guidelines. The proposed approach is validated through simulation studies and offers actionable strategies for generalizing findings from common yet constrained settings—such as using students as proxies for professional developers—to broader target populations. By explicitly modeling the mechanisms underlying population differences, the method significantly enhances the practical applicability and reliability of experimental results across diverse real-world contexts.
To address latent bias arising from unmeasured confounding in observational studies, this paper proposes a novel causal inference paradigm based on sample splitting: data are partitioned into planning and analysis samples, where the former adaptively selects robust design parameters (e.g., matching strategies, covariate sets), and the latter yields unbiased causal estimates. The method innovatively integrates multiple testing correction, heteroskedasticity-robust covariance estimation, and formal sensitivity analysis—extending support to multiple outcome variables, thereby relaxing the conventional single-outcome assumption. We establish theoretical guarantees of statistical validity under latent bias. Simulation studies demonstrate substantially higher statistical power than benchmark methods under strong unmeasured confounding. Empirical application to assessing the multidimensional impacts of floods in Bangladesh confirms practical feasibility and robustness.
This paper addresses the concurrent presence of three systematic biases in causal inference: interference (where an individual’s treatment affects others’ outcomes), unmeasured confounding, and lack of transportability across populations. We propose the first unified weighted sensitivity analysis framework that jointly quantifies the impact of all three biases on causal effect estimation. Our approach introduces interpretable sensitivity parameters and employs a weighting-based estimation strategy that accommodates unmeasured confounding while explicitly modeling interference structures and constraints on cross-population extrapolation. Empirical evaluations across multiple real-world settings demonstrate that the method robustly assesses bias magnitude and enhances the credibility of causal estimates. It provides an interpretable, scalable tool for causal inference in complex, dependent environments—such as social networks and public health interventions—where traditional assumptions of independence and identifiability fail.
This study addresses the challenge of ensuring rigor in causal inference under multi-source heterogeneous data fusion by proposing a structured design paradigm grounded in the target trial framework. The approach explicitly incorporates the target population and its sampling model into the causal analysis, systematically integrating external controls, generalizability, and transportability assessments through data element alignment, transparent assumption articulation, and emulation of the target trial. Its key innovation lies in anchoring the entire framework to a precise definition of the target population, thereby identifying and mitigating irreconcilable conflicts across data sources. This strategy enhances both the reliability and interpretability of causal conclusions derived from complex, real-world data ecosystems.
研究解决了跨网络干扰下的因果效应迁移问题,通过扩展选择图并提出TranCE算法,结合干预结果模型、域密度比校正和交叉拟合推断方法。
This study addresses the challenge of constructing high-quality external control arms for target populations using multi-source data to reliably estimate causal effects of alternative treatments. To this end, the authors propose two identification strategies—one based on the transportability of potential outcome means and the other on the transportability of effect measures—and develop a semiparametric, doubly robust augmented weighting estimator. This estimator integrates models for trial participation probability, treatment assignment probability, and conditional outcome mean, maintaining robustness under partial model misspecification. Theoretical analysis establishes its asymptotic efficiency, while simulation studies demonstrate superior finite-sample performance compared to methods relying solely on modeling or weighting. The approach is successfully applied to the ACCEPT and PHOENIX 1 trials, yielding reliable assessments of efficacy differences among biologic therapies for psoriasis.
This study addresses the challenge of efficiently estimating causal effects under confounding when experimental budgets are limited. The authors propose a novel approach that integrates instrumental variable regression with Gaussian graphical models, leveraging prior knowledge of partial joint distributions to optimize the allocation between fully observed samples and partially observed data (e.g., only \(X_{12}\)). Under a fixed budget constraint, this method analytically derives the optimal sampling scheme that minimizes the asymptotic variance of the causal effect estimator—a solution not previously available in closed form. Theoretical analysis demonstrates that the proposed allocation significantly reduces both the total budget and the number of complete observations required to detect non-zero causal effects. Empirical validation in automotive analytics and drug discovery underscores the method’s practical utility alongside its theoretical contributions.