Wasserstein Causal Forests for Distribution-Valued Outcomes

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of traditional causal inference, which is confined to mean effects and struggles to capture how interventions alter the shape of outcome distributions. To overcome this, we propose the Wasserstein-based Causal Forest (WCF) method. WCF defines distributional average and conditional treatment effects, precisely quantifying distributional shifts and deformations by incorporating reference distance comparisons and finite grid transformation techniques. Experiments demonstrate that WCF achieves superior estimation accuracy across most scenarios. Furthermore, an empirical application to the STAR dataset reveals that small class sizes significantly improve scores among high-performing students while exacerbating distributional inequality. These findings validate the effectiveness of the proposed approach in extending beyond conventional mean-level analysis.
📝 Abstract
This paper proposes Wasserstein Causal Forests (WCF) for settings in which each unit's outcome is itself a probability distribution. This study also defines finite-grid transformed average and conditional average treatment effects, including a reference-distance contrast that asks whether treatment moves unit-level distributions toward a prespecified benchmark. Simulations cover null effects, location and shape changes, limited overlap, equal-mean but different laws, heterogeneous effects, multimodality, and structural zeros. WCF is most accurate on the conditional-law metric in most reported designs and sharply improves reference-effect estimation in the principal location-and-shape settings, but it is less accurate than the forest baselines for multimodal settings. WCF is applied to the famous Project STAR \citep{word1990state}, revealing that small classes alter more than the mean: they raise within-grade mathematics achievement by $0.161$ standard deviations on average (SE $0.028$); but the gain is not a uniform location shift, it is larger in the upper part of the classroom score distribution ($+0.179$ at the ninetieth percentile versus $+0.091$ at the tenth) and, in descriptive stratum estimates, largest in the schools serving the most economically disadvantaged students (highest free-lunch quartile, $+0.262$, versus $+0.092$ to $+0.164$ elsewhere), while overall dispersion is essentially unchanged.
Problem

Research questions and friction points this paper is trying to address.

Causal inference
Distribution-valued outcomes
Wasserstein distance
Treatment effects
Heterogeneous effects
Innovation

Methods, ideas, or system contributions that make the work stand out.

Wasserstein Causal Forests
Distribution-Valued Outcomes
Conditional Average Treatment Effect
Reference-Distance Contrast
Heterogeneous Treatment Effects
🔎 Similar Papers
💼 Related Jobs
No related jobs found.