🤖 AI Summary
This study addresses the challenge in counterfactual prediction where observational data often fail to distinguish among latent causal graphs, leading to biased estimates when relying on a single causal structure. To overcome this, we propose CAFE, a framework that introduces amortized inference to this task for the first time. By pre-training a Transformer to approximate the Bayesian counterfactual posterior distribution, CAFE integrates multiple potential causal models through a single forward pass, enabling highly efficient inference. This approach effectively resolves causal structure uncertainty and yields precise individual-level outcome predictions under identifiable scenarios. Furthermore, CAFE demonstrates robust performance surpassing traditional assumption-bound limitations in real-world applications across manufacturing and agriculture, establishing it as a practical solution for counterfactual reasoning under structural ambiguity.
📝 Abstract
Counterfactual prediction estimates an individual's outcome under an alternative intervention given their factual observations. Such outcomes are generally not identifiable from observational data without additional assumptions. Even within the class of fully observed additive noise models (ANMs), different causal graphs can generate the same observational distribution yet imply different individual counterfactual outcomes. Predictions based on a single estimated graph ignore this structural uncertainty. We therefore target a Bayesian counterfactual posterior predictive distribution that combines predictions from plausible SCMs. We introduce CAFE (\textbf{C}ounterf\textbf{A}ctual Prediction via \textbf{F}ast Posterior \textbf{E}stimation), an amortized inference framework that directly approximates the Bayesian counterfactual posterior predictive distribution. We pretrain a transformer-based model on synthetic counterfactual tasks generated from a diverse prior over ANMs. Given an observational dataset, an individual's factual observations, and an intervention, CAFE approximates the corresponding posterior predictive distribution in a single forward pass. Experiments show that CAFE accurately predicts individual counterfactual outcomes in identifiable settings and approximates the posterior predictive distribution when structural uncertainty induced by observationally indistinguishable causal graphs exists. Strong performance in realistic manufacturing and viticulture settings further demonstrates its empirical robustness beyond the assumptions of the training prior.