Active Inference as a Convex Markov Decision Process

πŸ“… 2026-07-22
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work integrates the minimization of expected free energy from active inference into an optimizable decision-making framework by formally casting it as a Convex Markov Decision Process (Convex MDP) for the first time, unifying epistemic exploration and pragmatic goals in the space of state marginal distributions. By revealing that expected free energy corresponds to a policy-dependent instrumental reward, the study establishes compatibility with dynamic programming and actor-critic methods. Leveraging convex optimization and mirror descent, it derives policy optimization algorithms applicable to finite-horizon, discounted, and average-reward settings. This approach provides theoretical guarantees for policy improvement and bridges active inference with modern reinforcement learning through a rigorous theoretical foundation.
πŸ“ Abstract
Active Inference (AIF) frames adaptive behavior as the minimization of expected free energy (EFE), combining epistemic and pragmatic objectives within a single variational principle. We frame AIF as policy optimization and show that, for closed-loop control policies, EFE minimization can be formulated as a convex Markov decision process (MDP). In this formulation, the pragmatic terms are linear in the predictive state marginals and therefore equivalent to reward maximization in a latent MDP, while the epistemic value introduces a nonlinear component that distinguishes EFE minimization from standard reinforcement learning. This perspective further reveals the epistemic drive of active inference as a policy-dependent (performative) reward. We analyze finite-horizon, discounted, and average-reward formulations of EFE and derive a mirror descent (MD) algorithm that locally linearizes the objective around the current state marginals, yielding a policy-dependent reward that is compatible with actor-critic methods and dynamic programming. Finally, we argue that coupling world-model learning with policy optimization gives active inference the structure of performative reinforcement learning, providing a route toward grounding active inference within modern reinforcement learning and optimization theory, including convergence analysis and principled policy improvement guarantees.
Problem

Research questions and friction points this paper is trying to address.

Active Inference
Expected Free Energy
Convex MDP
Epistemic Value
Performative Reinforcement Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Active Inference
Convex MDP
Expected Free Energy
Performative Reinforcement Learning
Mirror Descent
πŸ”Ž Similar Papers