Score
Identifying and estimating causal effects that vary across subpopulations by accounting for confounding and interaction with covariates, and quantifying how treatment impacts differ in magnitude or persistence across groups or segments.
Estimating causal effects of multiple air pollutants on multiple health outcomes under unmeasured confounding remains a fundamental challenge in environmental epidemiology. Method: We propose the first joint partial identification framework tailored to the multi-treatment–multi-outcome setting. Leveraging the factor confounding assumption to model residual dependence, we introduce novel joint constraints across multiple estimands—tightening individual effect bounds—and establish conditions under which negative control variables enable point identification. Our method integrates factor modeling, partial identification set optimization, and robust numerical algorithms. Results: Empirical analysis on Medicare claims data demonstrates that estimated effects of pollutants—including PM₂.₅, NO₂, and O₃—on cardiovascular and respiratory outcomes exhibit robustness to unmeasured confounding. This work advances causal inference in environmental health by providing a principled, computationally tractable framework for bounding heterogeneous treatment effects in high-dimensional, confounded settings.
Exposure contrasts—commonly used in causal inference to estimate treatment and spillover effects—may yield estimates with signs opposite to the true unit-level causal effects, particularly under interference. Method: We systematically characterize the causal interpretability boundary of exposure contrasts within a nonparametric framework, formalizing exposure mappings via causal graphs. We derive verifiable necessary and sufficient conditions for sign consistency, ensuring that exposure contrasts robustly reflect the direction of causal effects—even under arbitrary assignment mechanisms (e.g., cluster-randomized trials, network experiments, or observational data with peer selection) and arbitrary interference structures. Contribution/Results: This work establishes, for the first time, a function-form–free theory guaranteeing sign preservation of exposure contrasts. It provides a rigorous foundation for reliable identification of spillover effects in social network analysis and policy evaluation, bridging theoretical causality and empirical practice without restrictive modeling assumptions.
This study addresses the problem of extrapolating causal effects from multi-site randomized controlled trials (RCTs) to a new target site with baseline survey data only. To handle site-level population heterogeneity and unobserved confounding, we propose modeling baseline covariates as functional data—thereby capturing site-specific confounding structures—for the first time. We then develop a design-oriented, nonparametric method to construct an optimal finite-dimensional feature space, ensuring optimal convergence rates for conditional average treatment effect (CATE) estimation. Our approach integrates functional data analysis, nonparametric regression, and causal transfer learning theory. Evaluated across five integrated multi-site RCTs on cash transfer programs, the method significantly improves prediction accuracy of treatment effects at target sites and quantifies the estimation gain attributable to adaptive transfer.
This study addresses the inconsistency in causal effect estimates between observational studies and randomized controlled trials (RCTs) by proposing the first unified framework for decomposing causal effect heterogeneity. The framework systematically identifies and quantifies three sources of heterogeneity: differences in covariate distributions, variation in mediating pathways, and shifts in outcome-generating mechanisms. Methodologically, it formally defines effect decomposition across data types (observational vs. experimental), integrating causal inference, sensitivity analysis, and decomposition modeling, while enabling robust parameter estimation under multiple hypotheses. Evaluated through simulation studies and an empirical analysis of the “Moving to Opportunity” experiment, the framework demonstrates improved interpretability, robustness, and policy generalizability in synthesizing evidence from heterogeneous data sources.
This study addresses selection bias in estimating long-term causal effects—such as graduation rates—from observational studies. We propose a novel control function approach that leverages experimental estimates of treatment effects on short-term outcomes (e.g., eighth-grade test scores) to correct for unobserved confounding in large-scale administrative observational data. Our method integrates insights from difference-in-differences estimation, covariate balancing, and cross-sample effect calibration, enabling the first systematic correction based on heterogeneity in short-term treatment effects. By bridging randomized experiments and observational datasets, the framework jointly preserves internal validity from experiments and external representativeness from administrative records, overcoming inferential limitations inherent to single-data-source designs. Empirical validation using the STAR randomized experiment and New York State school administrative data demonstrates substantial improvements in both accuracy and external validity of estimated causal effects of class size on academic performance.
This study addresses the challenge of identifying heterogeneous treatment effects while controlling the false discovery rate (FDR) in matched observational studies, where limited sample sizes and unmeasured confounding pose significant obstacles. The authors propose a novel method that adaptively discovers interpretable subgroups defined by covariate thresholds under many-to-one matching designs. Their approach achieves exact FDR control at the subgroup level for the first time and integrates sensitivity analysis models to account for unobserved confounding, leveraging multiple controls to enhance statistical power. Theoretical analysis, simulations, and empirical evaluation demonstrate that the method outperforms existing baselines in both accuracy and power when estimating heterogeneous economic returns to college education.
This study addresses the limitations of traditional control-based causal inference methods—such as matching and difference-in-differences—in settings characterized by pervasive or structurally ambiguous spillover effects, where reliance on uncontaminated control units impedes accurate identification of both average direct and spillover effects. Within the potential outcomes framework, this work provides the first systematic comparison between control-based and prediction-based counterfactual approaches—including interrupted time series and machine learning control—in terms of their identification capabilities. Through simulation and empirical analyses, the authors demonstrate that in environments with widespread interference, prediction-based methods can more reliably estimate certain causal parameters over short horizons, circumventing the stringent assumption of unperturbed units and thereby offering a promising alternative for causal inference under complex interference.
This study investigates the validity conditions for identifying causal effects using changes in treatment variables rather than their levels, and examines the relationship between this approach and conventional methods. By developing two non-nested structural models and integrating structural causal modeling with difference-in-differences and two-period fixed-effects regression, the authors theoretically demonstrate that strategies based on treatment changes and treatment levels are generally non-nested but become equivalent under specific conditions. They propose a corresponding overidentification test to assess these conditions. Simulation evidence confirms the favorable finite-sample performance of the proposed method, and an empirical application to cigarette demand estimation supports its practical validity. The work clarifies the fundamental distinctions and connections between these two causal identification strategies, thereby extending the methodological foundations of causal inference.
In observational causal inference, accurate identification of confounding variables is critical for reliable causal estimates. This work proposes ConfoundingSHAP, a novel method that uniquely adapts Shapley values to quantify the confounding strength of covariates. By formulating a Shapley game specifically tailored to confounding effects and designing an appropriate value function, the approach precisely measures each variable’s contribution to bias in causal estimation. Furthermore, it integrates TabPFN to enable efficient and scalable evaluation of adjustment sets without repeated model retraining. Empirical results across multiple datasets demonstrate that ConfoundingSHAP accurately identifies key confounders and provides interpretable, trustworthy insights into the sources of confounding.
This study addresses the limitations of conventional difference-in-differences methods in analyzing heterogeneous treatment effects across groups, which are often confounded by differences in covariate distributions, conservative inference procedures, and restrictive parametric interaction structures. To overcome these challenges, the paper proposes a novel estimator—the Balanced Group Average Treatment Effect on the Treated (BGATT)—which, under the parallel trends assumption, effectively disentangles covariate composition differences from genuine treatment effect heterogeneity. BGATT offers clear identifiability and interpretability while accommodating flexible, high-dimensional modeling. The authors construct an influence-function-based estimator that achieves √n-consistency and asymptotic normality, enabling efficient nuisance parameter estimation via machine learning. Both theoretical analysis and simulation studies demonstrate that the proposed method delivers superior finite-sample performance, substantially enhancing the accuracy and robustness of heterogeneity assessments.