Score
Designs and assesses methods that estimate the conditional average treatment effect (CATE) — the expected difference in outcomes between treatment and control conditional on covariates — by building predictive models and uncertainty quantification for heterogeneous effects. This includes constructing and validating CATE estimators, developing procedures to extrapolate or transport effects to different populations, and producing interpretable visualizations or summaries of learned heterogeneity or policy-mixing strategies.
Conventional conditional average treatment effect (CATE)-based methods struggle to identify treatment effect heterogeneity when effect modifiers are unobserved or subject to severe measurement error. Method: We propose a variance-comparison inference framework that does not require fully observed covariates. Leveraging the variance difference of potential outcomes as a novel causal identification anchor, we construct a doubly robust and asymptotically linear nonparametric estimator, integrating causal machine learning with a variance-sensitive testing paradigm. Contribution/Results: We establish theoretical consistency and asymptotic normality under weak regularity conditions. In a reanalysis of a randomized controlled trial, our method detects statistically significant heterogeneity in therapeutic hypothermia efficacy. The approach demonstrates robustness across diverse data-generating mechanisms and overcomes the strong reliance of CATE-based methods on high-fidelity covariate measurement.
This work addresses the high sample complexity of traditional conditional average treatment effect (CATE) methods, which require large datasets to enable precise interventions in heterogeneous populations. The authors propose a novel framework that achieves near-optimal aggregate intervention outcomes with only \(O(M/\varepsilon)\) samples by leveraging coarse treatment effect estimates combined with a greedy allocation strategy under budget constraints. The key insight lies in distinguishing between the goals of effect estimation and intervention allocation, demonstrating that highly accurate CATE estimates are unnecessary for effective decision-making. By exploiting structural properties of the underlying data distribution and incorporating stratified sampling, the method further reduces sample requirements. Empirical evaluations on multiple real-world randomized controlled trial datasets confirm the approach’s efficiency and practical utility.
This paper addresses the bias in conditional average treatment effect (CATE) estimation under hidden confounding. We propose a calibration method that does not require access to covariates from randomized controlled trials (RCTs), instead leveraging a small RCT dataset to calibrate potential outcomes in large-scale observational data. Our core innovation is a pseudo-confounder generator that jointly performs adversarial distribution alignment and deep CATE estimation, enabling implicit alignment between the potential outcome spaces of observational and RCT data—even without RCT covariates—thereby relaxing the conventional conditional ignorability assumption. By integrating causal inference with representation learning, our approach significantly reduces CATE estimation error on both synthetic and real-world healthcare datasets. It is especially suitable for privacy-sensitive settings where RCT covariates are unavailable or restricted. Empirical results demonstrate superior robustness and accuracy compared to state-of-the-art methods.
This paper addresses the challenge of quantifying individual-level treatment risk in binary-outcome settings. We propose the Fraction of Negative Average (FNA) metric—the proportion of individuals whose outcomes deteriorate upon treatment—thereby complementing the Conditional Average Treatment Effect (CATE), which captures only subgroup-level averages and obscures individual harm. Under the ignorability assumption, we introduce the Pearson correlation coefficient between potential outcomes as a sensitivity parameter and derive tight, feasible theoretical bounds for FNA—substantially improving upon the classical Fréchet–Hoeffding bounds. We establish an analytical relationship among FNA, CATE, and the correlation coefficient, revealing the counterintuitive phenomenon that positive CATE can coexist with substantial individual harm. We further propose principled guidelines for selecting plausible correlation ranges and develop a nonparametric estimator for FNA that is consistent and asymptotically normal.
In precision medicine, existing conditional average treatment effect (CATE) estimation methods struggle to identify key covariates driving treatment effect heterogeneity. To address this, we propose Treatment Effect Variable Importance Measures (TE-VIMs), the first nonparametric generalization of ANOVA for causal heterogeneity analysis—model-agnostic and statistically valid. Leveraging the efficient influence curve (EIC), we construct asymptotically optimal estimators that unify nonparametric causal inference, machine learning, and semiparametric efficiency theory. Extensive simulations and real-world clinical studies demonstrate that TE-VIMs substantially improve identification accuracy of critical effect modifiers, enhance interpretability and clinical utility of CATE models, and bridge the interpretability gap between policy learning and CATE estimation.
Traditional average treatment effects fail to capture individual heterogeneity, and under high-dimensional settings, the sublevel set structure of the conditional average treatment effect (CATE) function is complex, lacking a concise global measure of heterogeneity. This work formalizes the probability curve of CATE sublevel sets as a target parameter for the first time, revealing its non-pathwise differentiability. By integrating Grenander-type monotone estimation with debiased machine learning techniques, the authors develop a nonparametric inference framework. The proposed estimator demonstrates strong finite-sample performance and is applied empirically to randomized trial data on diabetes medication, effectively uncovering heterogeneous treatment effects across subpopulations.
Reliably evaluating the goodness-of-fit of conditional average treatment effect (CATE) estimates derived from observational data remains a critical challenge for applying causal inference in policy and personalized decision-making. This work proposes the CAFE framework, which introduces the first validation approach directly targeting CATE estimation—rather than the full outcome model—by leveraging auxiliary randomized controlled trial (RCT) data. CAFE stratifies the covariate space using propensity scores and conducts hypothesis tests based on group-level treatment effect comparisons. To enhance sensitivity to local model misspecification, it incorporates a maximal test statistic and employs a two-stage procedure to detect potential unmeasured confounding. The framework is compatible with both parametric models and flexible machine learning methods such as causal forests. Extensive experiments demonstrate that CAFE effectively identifies CATE model mismatches, offering a reliable assessment when both RCT and observational data are available.
This work addresses the challenge of obtaining well-calibrated uncertainty intervals for conditional average treatment effects (CATE) in “few-treated” settings, where the number of treated units is substantially smaller than that of control units. Existing methods, such as the X-Learner, often fail to properly quantify uncertainty under such imbalance. The authors propose GP-CATE, a Bayesian approach that jointly models the outcome surfaces of both treatment and control groups using Gaussian processes. By directly incorporating the heightened uncertainty associated with the small treated group into the posterior inference, GP-CATE avoids the model perturbation bias inherent in conventional two-stage estimators. To the best of the authors’ knowledge, this is the first method to deliver calibrated uncertainty estimates for CATE in few-treated scenarios. Empirical evaluations on synthetic and semi-synthetic datasets demonstrate that GP-CATE consistently outperforms benchmark methods—including X-Learner, Causal Forest, and BART—producing confidence intervals with accurate coverage and reasonable width, even under severe data scarcity.
This study addresses the critical need for accurate estimation of the conditional average treatment effect (CATE) for specific events in survival analysis under competing risks and right censoring, which is essential for personalized medicine. The authors propose a meta-learner–based framework in a binary treatment setting, defining CATE as the absolute risk difference at a fixed time point. They systematically evaluate six meta-learner strategies that combine either Cox regression or random survival forests to model event-specific risks, paired with elastic net or random forest models to directly estimate CATE. Comprehensive simulations encompassing diverse risk structures, treatment effect heterogeneity, treatment assignment mechanisms, and censoring levels are conducted to assess performance. Based on empirical findings, the study offers practical modeling recommendations and releases the open-source R package crsurvlearners to facilitate broader application.
This study addresses the limited external validity of randomized controlled trials, which often enroll participants systematically different from the target population—particularly when effect modifiers are unevenly distributed—rendering the average treatment effect (ATE) inadequate for capturing treatment effect heterogeneity. Within a nested trial framework, the authors propose a novel approach to unbiasedly generalize the conditional average treatment effect (CATE) to the entire eligible population based on pre-specified effect modifiers. Leveraging semiparametric theory and data-adaptive estimation, the method constructs pseudo-outcomes via conditional influence functions and employs local linear kernel regression with cross-fitting to mitigate overfitting. Simulations and an empirical application to the Coronary Artery Surgery Study (CASS) demonstrate that the proposed estimator robustly recovers and enables valid inference on heterogeneous treatment effects in the target population.