Score
Designing and running targeted experiments or interventions that isolate causal effects by controlling confounds and measuring outcomes, and validating whether internal model or learned representations encode mechanistic structure beyond surface inputs. This covers field-study design, perturbation tests, and validation strategies to attribute observed effects to the intended manipulations.
When randomized controlled trials are infeasible—as in rare diseases or oncology—effectively leveraging external data becomes a critical challenge. This work proposes a six-step causal inference–based scientific framework that systematically integrates external control arm designs from single-arm and hybrid control trials, unifying Bayesian dynamic borrowing, frequentist approaches, and modeling of covariate shift and outcome drift. Centered on causal identifiability, the framework clarifies, for the first time, the trade-off between efficiency and robustness inherent in external control methodologies and underscores the necessity of sensitivity analyses in regulatory decision-making. Through a systematic literature review and empirical evaluation, the study provides a coherent guide and accompanying software tools to support both theoretical integration and practical application of external data.
This paper addresses the challenge of estimating causal effects when the target variable cannot be directly intervened upon and the underlying mechanism is complex—nonlinear, high-dimensional, and confounded. We propose the first active experimental design framework tailored for *indirect experiments*. Methodologically, we formulate a bilevel optimization model that integrates kernel-based estimation with adaptive sequential experimental design, yielding an analytically tractable and computationally efficient estimator for upper and lower bounds on the causal effect. Our key contributions are: (1) the first systematic formalization of feasibility conditions for indirect intervention under nonlinear confounding; and (2) dynamic narrowing of the causal bound gap to precisely localize the target query value. Extensive synthetic experiments across diverse settings demonstrate that our method significantly improves causal effect identification accuracy, with faster convergence of bound width compared to state-of-the-art baselines.
To address latent bias arising from unmeasured confounding in observational studies, this paper proposes a novel causal inference paradigm based on sample splitting: data are partitioned into planning and analysis samples, where the former adaptively selects robust design parameters (e.g., matching strategies, covariate sets), and the latter yields unbiased causal estimates. The method innovatively integrates multiple testing correction, heteroskedasticity-robust covariance estimation, and formal sensitivity analysis—extending support to multiple outcome variables, thereby relaxing the conventional single-outcome assumption. We establish theoretical guarantees of statistical validity under latent bias. Simulation studies demonstrate substantially higher statistical power than benchmark methods under strong unmeasured confounding. Empirical application to assessing the multidimensional impacts of floods in Bangladesh confirms practical feasibility and robustness.
In high-dimensional observational settings of randomized controlled trials (RCTs), causal treatment effect estimation is prone to modeling and sampling selection biases. Method: We introduce ISTAnt—the first real-world visual causal benchmark grounded in ant-behavior RCTs—and theoretically prove that classification accuracy cannot serve as a proxy for causal estimation quality. We further propose representation-learning principles tailored for scientific causal inference. Contribution/Results: Through tripartite validation—rigorous theoretical analysis, controlled synthetic experiments, and real biological experiments—across 6,480 large-scale fine-tuned models built upon state-of-the-art vision backbones, we demonstrate that common deep learning practices (e.g., loss design, data sampling) induce substantial systematic biases, and classification performance exhibits no strong correlation with causal estimation accuracy. Our work establishes a reproducible benchmark, theoretical criteria, and practical guidelines for high-dimensional causal inference.
This study addresses the challenge of integrating randomized or single-arm clinical trials with external experimental or observational data to enable cross-study treatment comparisons and improve estimation precision of treatment effects. Methodologically, building upon the potential outcomes framework, we first develop a unified identification strategy for hybrid-data designs, systematically characterizing identifiability conditions across diverse designs—including historical controls, synthetic controls, and anchoring estimators—and propose a generalizable taxonomy of such designs along with corresponding causal inference principles. Our contribution lies in filling a critical theoretical gap in regulatory science regarding the rigorous integration of external controls, thereby establishing a methodological foundation for leveraging real-world evidence to complement trial-based evidence in pharmaceutical and medical device evaluation. This advancement significantly enhances the transportability of evidence and its applicability to regulatory decision-making.
This paper addresses the challenge of confounder selection in observational studies by proposing an interactive, iterative method that requires neither a pre-specified causal graph nor a complete set of candidate variables. Grounded in latent projection theory, the method dynamically expands a causal graph through successive user-provided local adjustment sets and automatically identifies a minimal “principal adjustment set,” thereby determining whether confounding is controllable. Its key contributions are threefold: (1) it is the first approach to achieve sound and complete confounding control assessment without prior structural assumptions on the causal graph; (2) it makes no assumptions about causal relationships among potential confounders; and (3) it bridges theoretical rigor with practical feasibility. Both theoretical analysis and empirical evaluation demonstrate that, under correct user feedback, the algorithm accurately identifies admissible adjustment sets and correctly determines confounding controllability.
This study addresses the limitations of traditional control-based causal inference methods—such as matching and difference-in-differences—in settings characterized by pervasive or structurally ambiguous spillover effects, where reliance on uncontaminated control units impedes accurate identification of both average direct and spillover effects. Within the potential outcomes framework, this work provides the first systematic comparison between control-based and prediction-based counterfactual approaches—including interrupted time series and machine learning control—in terms of their identification capabilities. Through simulation and empirical analyses, the authors demonstrate that in environments with widespread interference, prediction-based methods can more reliably estimate certain causal parameters over short horizons, circumventing the stringent assumption of unperturbed units and thereby offering a promising alternative for causal inference under complex interference.
This study addresses the challenge of identifying heterogeneous treatment effects while controlling the false discovery rate (FDR) in matched observational studies, where limited sample sizes and unmeasured confounding pose significant obstacles. The authors propose a novel method that adaptively discovers interpretable subgroups defined by covariate thresholds under many-to-one matching designs. Their approach achieves exact FDR control at the subgroup level for the first time and integrates sensitivity analysis models to account for unobserved confounding, leveraging multiple controls to enhance statistical power. Theoretical analysis, simulations, and empirical evaluation demonstrate that the method outperforms existing baselines in both accuracy and power when estimating heterogeneous economic returns to college education.
This study addresses the problem of testing whether a treatment effect operates entirely through observed mediators and identifying causal mechanisms under control for covariates. The authors propose a statistical test based on double machine learning, extending— for the first time—the joint evaluation of full mediation and causal mechanism identification to non-randomized treatment settings. By integrating conditional independence testing, the method achieves root-n consistent and asymptotically normal inference even in the presence of high-dimensional covariates. Simulation studies demonstrate favorable finite-sample performance, and the approach is successfully applied to two randomized experiments examining maternal mental health and social norms.
This study addresses the failure of conventional causal inference methods in group interaction experiments, where within-group interactions and interference effects violate standard assumptions. The authors develop a design-based causal inference framework that systematically characterizes identifiability under various scenarios—such as fixed or random group assignment and presence or absence of interference—and proposes corresponding inference strategies. Innovatively, they introduce a coupling strategy to handle complex dependence structures, integrating sparse-sampling asymptotics, cluster-robust inference, and the potential outcomes framework. They demonstrate that, even under interference, cluster-robust methods consistently estimate marginalized exposure effects. Moreover, when interference is absent and assignment is randomized, the framework naturally reduces to the standard individual-level randomized experiment, thereby preserving compatibility with classical individual-level inference.
This study addresses the challenge of identifying direct, indirect, and total effects in spatial interference experiments, where spillover effects—particularly under cluster randomization—render the level of the spillover kernel non-identifiable. The authors develop a unified causal framework that expresses estimands as linear functionals of the spillover kernel and introduce an “anchoring assumption” to resolve the level identification problem. By positing separability between the kernel’s shape and its level, they demonstrate that all estimators inherently depend on the choice of anchor and propose a “level leverage” measure to quantify sensitivity to the unidentifiable level. Integrating linear exposure mappings, bias decomposition, and design augmentation strategies, the framework leverages prior geometric structure to compute leakage terms and shape errors. The work further shows that conventional cluster-based analyses are special cases of implicit anchoring, thereby unifying existing approaches and offering principled criteria for deciding whether to augment the experimental design or select an anchor.