🤖 AI Summary
This study addresses the challenge of effectively balancing intrinsic motivation-driven exploration and goal-directed decision-making for agents operating in partially observable environments. To this end, it extends the maximum occupancy principle to partially observable settings by introducing a belief-based reasoning mechanism. Furthermore, this work proposes a Bellman reformulation of expected free energy, enabling offline value iteration solutions across the entire belief state space. This methodological advance significantly improves algorithmic tractability under uncertainty. Experimental results demonstrate that the proposed agent dynamically switches between exploration and exploitation based on energy and belief states, thereby overcoming the limitation of conventional active inference strategies that tend to converge on single sources. Ultimately, the approach achieves superior adaptive behavior generation.
📝 Abstract
Intrinsic motivation plays a central role in adaptive and goal-directed behavior by conferring agents reward-independent objectives and biases useful to act in noisy and uncertain environments. Active Inference addresses the problem of acting in a partially observable environment through a principled framework for belief updating and action selection. A key component of Active Inference is the specification of prior preferences, which shapes behavior by encoding desirable future outcomes. An intrinsic motivation approach called the Maximum Occupancy Principle (MOP) proposes that agents act so as to maximize occupancy over future paths of states and actions, with no preferences or epistemic targets. Despite its simple formulation, MOP gives rise to rich and adaptive behaviors that combine exploratory variability with goal-directed dynamics. In this work, we extend MOP to partially observable environments and introduce a Bellman reformulation of the Expected Free Energy for Active Inference, both incorporating belief-based inference over hidden states as part of the agent state. The Bellman formulation enables tractable offline computation via value iteration over the full belief-state space. We compare the resulting behaviors in a set of minimal experimental settings with uncertain food sources. We find that MOP agents switch between goal-directed (food seeking) behavior and exploration between different food sources, depending on their energy available and their belief state. In contrast, Active Inference agents mostly inhabit regions around a single food source, a strategy having both high pragmatic and epistemic value. We finally compare with Empowerment, which is shown to be qualitatively similar to Active Inference.