🤖 AI Summary
Estimating causal effects from real-world data is challenging due to hidden confounding and spatiotemporal interference. This work proposes the first proximal causal inference framework tailored for spatiotemporal settings. It jointly models local and neighborhood confounding through treatment- and outcome-induced proxy variables and constructs a confounding bridge function that satisfies both the exclusion restriction and spatiotemporal completeness conditions, enabling identification of potential outcomes without explicitly recovering latent confounders. The method integrates a Transformer encoder to learn proxy representations, a conditional mutual information discriminator to enforce the exclusion restriction, and a moment-matching network to solve for the bridge function, complemented by a stabilized weighting strategy to mitigate treatment support imbalance. Experiments demonstrate substantial improvements over baselines on synthetic data, offering the first theoretical identifiability guarantee for spatiotemporal causal effect estimation.
📝 Abstract
Estimating causal effects from real-world spatiotemporal data is challenging due to hidden confounders and interference. Standard causal identification methods assume conditional exchangeability given observed covariates, which fails whenever hidden confounders affect both treatment and outcomes - a common setting in domains such as climate, environmental policy, epidemiology, and regional economics. In this paper, we propose a novel spatiotemporal proximal causal inference framework that extends proximal identification theory to spatiotemporal settings. The proposed method jointly captures local and neighborhood-level confounding information by introducing treatment- and outcome-inducing proxies, and we derive a spatiotemporal outcome confounding bridge function that identifies the potential outcome without requiring direct recovery of the hidden confounder. We establish the identifiability of this bridge function under proxy exclusion restrictions and a spatiotemporal completeness condition, and show that the resulting estimator recovers the outcome through a proximal generalization of the g-computation formula. To operationalize this identification result, we propose a neural architecture that learns proxies via transformer-based spatiotemporal encoders - coupled with a conditional mutual information critic to enforce exclusion restrictions and a moment-matching network to guarantee that the learned bridge function satisfies the underlying identifying equation. We further introduce a stabilized weighting scheme to address treatment support imbalance. Experiments on synthetic datasets demonstrate that our approach achieves comparable performance to baseline causal inference methods, while providing, to our knowledge, the first theoretically grounded outcomes for the hidden confounding in the presence of spatiotemporal interference through a proximal causal inference framework.