🤖 AI Summary
This work addresses a key limitation in existing counterfactual reasoning approaches for sequential decision-making: their reliance on deterministic causal models, which fails to disentangle environmental stochasticity from unobserved confounding bias in Markov decision processes. To overcome this, the paper introduces the first counterfactual policy optimization framework grounded in probabilistic, non-deterministic causal models. The proposed method formally separates irreducible randomness from latent confounders and formulates a robust optimization problem amenable to sensitivity analysis. By integrating probabilistic causal modeling, counterfactual inference, and reinforcement learning–based policy optimization, the framework is validated on a sepsis treatment simulator, successfully identifying treatment policies robust to diabetes status—a global hidden confounder.
📝 Abstract
Counterfactual inference approaches for sequential decision-making typically assume deterministic causal models, where all randomness stems from latent variables. However, Markov Decision Processes (MDPs) are inherently stochastic. We address this by formalising counterfactual policy optimisation under probabilistic nondeterministic causal models, which properly separates latent confounding from irreducible stochasticity, and here propose a first practical optimisation problem for identifying robust counterfactual policies under a sensitivity analysis framework. We validate our approach on a sepsis treatment simulator, where diabetes status acts as a hidden global confounder.