đ€ AI Summary
This work addresses structured, partially observable multi-armed bandit problems characterized by complex environmental dependencies. To resolve the fundamental trade-off between over-exploration and convergence inherent in conventional approaches, we propose a general decision-making framework grounded in information maximization and the free energy principle. Our method introduces a hierarchical, structure-adaptive information measure, tailored to three canonical structured settings: graph-structured, hierarchical, and non-stationary environments. Integrating variational free energy inference, structured reward modeling, and sub-Gaussian/Gaussian-distribution-aware adaptation algorithms, we derive theoretically grounded optimal or near-optimal regret bounds. Empirical evaluations demonstrate that our approach significantly outperforms standard baselinesâincluding UCB and Thompson samplingâunder sparse feedback and non-stationary dynamics, achieving both computational efficiency and strong robustness across diverse structured domains.
đ Abstract
Information and free-energy maximization are physics principles that provide general rules for an agent to optimize actions in line with specific goals and policies. These principles are the building blocks for designing decision-making policies capable of efficient performance with only partial information. Notably, the information maximization principle has shown remarkable success in the classical bandit problem and has recently been shown to yield optimal algorithms for Gaussian and sub-Gaussian reward distributions. This article explores a broad extension of physics-based approaches to more complex and structured bandit problems. To this end, we cover three distinct types of bandit problems, where information maximization is adapted and leads to strong performance. Since the main challenge of information maximization lies in avoiding over-exploration, we highlight how information is tailored at various levels to mitigate this issue, paving the way for more efficient and robust decision-making strategies.