π€ AI Summary
This work integrates the minimization of expected free energy from active inference into an optimizable decision-making framework by formally casting it as a Convex Markov Decision Process (Convex MDP) for the first time, unifying epistemic exploration and pragmatic goals in the space of state marginal distributions. By revealing that expected free energy corresponds to a policy-dependent instrumental reward, the study establishes compatibility with dynamic programming and actor-critic methods. Leveraging convex optimization and mirror descent, it derives policy optimization algorithms applicable to finite-horizon, discounted, and average-reward settings. This approach provides theoretical guarantees for policy improvement and bridges active inference with modern reinforcement learning through a rigorous theoretical foundation.
π Abstract
Active Inference (AIF) frames adaptive behavior as the minimization of expected free energy (EFE), combining epistemic and pragmatic objectives within a single variational principle. We frame AIF as policy optimization and show that, for closed-loop control policies, EFE minimization can be formulated as a convex Markov decision process (MDP). In this formulation, the pragmatic terms are linear in the predictive state marginals and therefore equivalent to reward maximization in a latent MDP, while the epistemic value introduces a nonlinear component that distinguishes EFE minimization from standard reinforcement learning. This perspective further reveals the epistemic drive of active inference as a policy-dependent (performative) reward. We analyze finite-horizon, discounted, and average-reward formulations of EFE and derive a mirror descent (MD) algorithm that locally linearizes the objective around the current state marginals, yielding a policy-dependent reward that is compatible with actor-critic methods and dynamic programming. Finally, we argue that coupling world-model learning with policy optimization gives active inference the structure of performative reinforcement learning, providing a route toward grounding active inference within modern reinforcement learning and optimization theory, including convergence analysis and principled policy improvement guarantees.