Information maximization for a broad variety of multi-armed bandit games

📅 2025-03-20
📈 Citations: 0
✹ Influential: 0
📄 PDF
đŸ€– AI Summary
This work addresses structured, partially observable multi-armed bandit problems characterized by complex environmental dependencies. To resolve the fundamental trade-off between over-exploration and convergence inherent in conventional approaches, we propose a general decision-making framework grounded in information maximization and the free energy principle. Our method introduces a hierarchical, structure-adaptive information measure, tailored to three canonical structured settings: graph-structured, hierarchical, and non-stationary environments. Integrating variational free energy inference, structured reward modeling, and sub-Gaussian/Gaussian-distribution-aware adaptation algorithms, we derive theoretically grounded optimal or near-optimal regret bounds. Empirical evaluations demonstrate that our approach significantly outperforms standard baselines—including UCB and Thompson sampling—under sparse feedback and non-stationary dynamics, achieving both computational efficiency and strong robustness across diverse structured domains.

Technology Category

Machine Learning: Online Learning & BanditsReasoning under Uncertainty: Sequential Decision MakingSearch and Optimization: Learning to Search

Application Category

Economics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystemsGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Information and free-energy maximization are physics principles that provide general rules for an agent to optimize actions in line with specific goals and policies. These principles are the building blocks for designing decision-making policies capable of efficient performance with only partial information. Notably, the information maximization principle has shown remarkable success in the classical bandit problem and has recently been shown to yield optimal algorithms for Gaussian and sub-Gaussian reward distributions. This article explores a broad extension of physics-based approaches to more complex and structured bandit problems. To this end, we cover three distinct types of bandit problems, where information maximization is adapted and leads to strong performance. Since the main challenge of information maximization lies in avoiding over-exploration, we highlight how information is tailored at various levels to mitigate this issue, paving the way for more efficient and robust decision-making strategies.
Problem

Research questions and friction points this paper is trying to address.

Extends physics-based approaches to complex bandit problems
Adapts information maximization for efficient decision-making strategies
Addresses over-exploration by tailoring information at various levels
Innovation

Methods, ideas, or system contributions that make the work stand out.

Information maximization for multi-armed bandit games
Adaptation to complex structured bandit problems
Tailored information to prevent over-exploration
🔎 Similar Papers
No similar papers found.
đŸ’Œ Related Jobs
No related jobs found.
A
Alex Barbier––Chebbah
Institut Pasteur, UniversitĂ© Paris CitĂ©, CNRS UMR 3571, Decision and Bayesian Computation, 75015 Paris, France. ÉpimethĂ©e, Inria, Paris, France.
Christian L. Vestergaard
Christian L. Vestergaard
CNRS, Institut Pasteur
statistical physicscomputational biologydata analysisnetworksneuroscience
J
Jean-Baptiste Masson
Institut Pasteur, UniversitĂ© Paris CitĂ©, CNRS UMR 3571, Decision and Bayesian Computation, 75015 Paris, France. ÉpimethĂ©e, Inria, Paris, France.