MDP Planning as Policy Inference

📅 2026-02-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses planning in Markov decision processes (MDPs) over discrete domains by proposing a Bayesian inference framework that treats policies as latent variables. In this formulation, the unnormalized posterior probability of a policy increases monotonically with its expected return, thereby jointly capturing both optimality and uncertainty. Methodologically, the approach introduces policy consistency constraints and a cross-particle transition coupling mechanism, combined with variational sequential Monte Carlo (VSMC) and Thompson sampling to enable efficient posterior inference. Experiments in discrete environments—including grid worlds and Blackjack—demonstrate that the framework effectively reveals the underlying structure of policy distributions. The resulting behavior exhibits significant qualitative and statistical differences compared to Soft Actor-Critic, highlighting the advantages of explicitly modeling policy uncertainty.

Technology Category

Reasoning under Uncertainty: Sequential Decision MakingPlanning, Routing, and Scheduling: Planning with Markov Models (MDPs, POMDPs)Machine Learning: Calibration & Uncertainty Quantification

Application Category

Economics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystemsUser Modeling, Personalization and Recommendation: Psychology-informed user models and recommender systemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
We cast episodic Markov decision process (MDP) planning as Bayesian inference over _policies_. A policy is treated as the latent variable and is assigned an unnormalized probability of optimality that is monotone in its expected return, yielding a posterior distribution whose modes coincide with return-maximizing solutions while posterior dispersion represents uncertainty over optimal behavior. To approximate this posterior in discrete domains, we adapt variational sequential Monte Carlo (VSMC) to inference over deterministic policies under stochastic dynamics, introducing a sweep that enforces policy consistency across revisited states and couples transition randomness across particles to avoid confounding from simulator noise. Acting is performed by posterior predictive sampling, which induces a stochastic control policy through a Thompson-sampling interpretation rather than entropy regularization. Across grid worlds, Blackjack, Triangle Tireworld, and Academic Advising, we analyze the structure of inferred policy distributions and compare the resulting behavior to discrete Soft Actor-Critic, highlighting qualitative and statistical differences that arise from policy-level uncertainty.
Problem

Research questions and friction points this paper is trying to address.

MDP planning
policy inference
Bayesian inference
optimal behavior uncertainty
posterior distribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

policy inference
Bayesian planning
variational sequential Monte Carlo
Thompson sampling
policy uncertainty
🔎 Similar Papers
No similar papers found.