Calibrated Inference for the Conditional Average Treatment Effect in the Few-Placebo Regime via Gaussian Processes

📅 2026-05-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of obtaining well-calibrated uncertainty intervals for conditional average treatment effects (CATE) in “few-treated” settings, where the number of treated units is substantially smaller than that of control units. Existing methods, such as the X-Learner, often fail to properly quantify uncertainty under such imbalance. The authors propose GP-CATE, a Bayesian approach that jointly models the outcome surfaces of both treatment and control groups using Gaussian processes. By directly incorporating the heightened uncertainty associated with the small treated group into the posterior inference, GP-CATE avoids the model perturbation bias inherent in conventional two-stage estimators. To the best of the authors’ knowledge, this is the first method to deliver calibrated uncertainty estimates for CATE in few-treated scenarios. Empirical evaluations on synthetic and semi-synthetic datasets demonstrate that GP-CATE consistently outperforms benchmark methods—including X-Learner, Causal Forest, and BART—producing confidence intervals with accurate coverage and reasonable width, even under severe data scarcity.
📝 Abstract
Estimating how much an intervention helps a given individual the conditional average treatment effect (CATE) is increasingly central to decision-making in medicine, economics, and policy, where an estimate is most useful when accompanied by a calibrated uncertainty interval. We study the few-placebo regime, in which one treatment arm is much smaller than the other, as arises in unequal-allocation trials and small-holdout $A/B$ tests. The standard estimator in this setting is the X-Learner, and a natural way to obtain credible intervals is to make its second stage Bayesian. We show that these intervals under-cover: they contain the true effect less often than their nominal level. We trace this to a structural cause the X-Learner's regression target inherits the bias of a nuisance model fitted to the small arm, so the posterior is centered away from the true effect and we find that the standard remedy, regressing an orthogonal doubly-robust score, is also unreliable here, since the regime's limited overlap leaves the estimator either highly variable or, once stabilized, biased once more. Both consequences reflect a pattern that extends beyond causal inference: a separately estimated variance is attached to a point estimate of a hard-to-learn quantity, and the point estimate's bias is not captured by that variance. We propose GP-CATE, which models each arm's outcome surface with a Gaussian process, so the scarce arm's uncertainty enters the posterior directly rather than as an unmodelled bias. Across synthetic and semi-synthetic benchmarks, GP-CATE attains calibrated coverage where the estimators we compare against including Causal Forest and BART do not, at the cost of intervals that are appropriately wide when the data are uninformative.
Problem

Research questions and friction points this paper is trying to address.

Conditional Average Treatment Effect
Few-Placebo Regime
Uncertainty Calibration
Coverage
Causal Inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gaussian Processes
Conditional Average Treatment Effect
Uncertainty Calibration
Few-Placebo Regime
Causal Inference
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
E
Eichi Uehara
AFLO