Diverse Geometries, Frozen Weights: Robust Heterogeneous Treatment-Effect Estimation via Causal Expert Ensembles

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of determining appropriate inductive biases for estimating heterogeneous treatment effects from observational data characterized by overlap violations and imbalance. To this end, we propose the GeoACE framework, which introduces an outcome-independent, overlap-aware projection expert (O-Phi-ACE) that integrates anchor correction with complementary geometric diversity modeling. Robust aggregation is achieved through inverse doubly robust weighting coupled with a leakage-free frozen-weight ensemble strategy. Experimental evaluations demonstrate that the proposed method significantly reduces PEHE errors and achieves leading rankings across most benchmarks, thereby validating the effectiveness of both the diversified expert pool and the lossless aggregation strategy.
📝 Abstract
Estimating heterogeneous treatment effects from observational data is difficult because the most appropriate inductive bias varies with overlap, treatment imbalance, prognostic structure, and sample size. We introduce the Geometry-Diverse Anchor-Correction Expert Ensemble (GeoACE), a five-expert framework that combines a common anchor-correction estimator with complementary overlap-aware and outcome-guided geometries. Its task-level ensemble weights are learned only from internal validation predictions, frozen before test evaluation, and then applied to experts refitted on the complete development sample. The fifth expert, O-Phi-ACE, constructs an outcome-free, overlap-aware statistical projection from covariates and treatment assignment and replaces the anchor input with this lower-dimensional geometry. We evaluate GeoACE against 11 comparators on eight benchmark protocols. Adding O-Phi-ACE reduced mean sqrt(PEHE) relative to the four-expert ensemble on all seven benchmarks with individual-effect truth, winning 998 of 1,225 paired tasks; the change on JOBS policy risk was negligible. The five-expert ensemble ranked first on IHDP100, IHDPA, and IHDPB and second on NEWS, differing from the NEWS leader by 0.13%. Across the seven sqrt(PEHE) benchmarks it obtained the lowest observed average rank (3.714), although the omnibus Friedman and Iman-Davenport tests were not significant (p=0.328 and p=0.330). Using the same five frozen experts, inverse-DR weighting was consistently better than winner-take-all selection, convex DR fitting, R-stacking, and causal Q-aggregation in benchmark-balanced analyses, but was statistically indistinguishable from equal weighting and DR ridge shrinkage. The evidence therefore supports geometry-diverse expert libraries and leakage-free aggregation as a robustness strategy, not universal superiority of either GeoACE or one weighting rule.
Problem

Research questions and friction points this paper is trying to address.

Heterogeneous Treatment Effect Estimation
Observational Data
Inductive Bias
Causal Inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Heterogeneous Treatment Effect
Expert Ensemble
Anchor-Correction
Overlap-Aware Geometry
Frozen Weights
🔎 Similar Papers