🤖 AI Summary
This study addresses the insufficient exploration in traditional algorithms for non-stationary multi-armed bandits, which arises from focusing solely on current uncertainty. We propose the Future Information Directed Sampling (FIDS) algorithm, which extends the exploration objective from immediate rewards to predictive information structures for the first time. By explicitly exploring to acquire prior information about future optimal arms, FIDS overcomes the limitations of static exploration. Furthermore, it incorporates Bayesian inference and information-directed sampling, constructing an approximation framework based on offline supervised learning to enhance computational efficiency. Experiments on synthetic benchmarks demonstrate that FIDS effectively captures environmental dynamics, achieving cumulative regret close to the theoretical lower bound of Thompson Sampling.
📝 Abstract
Exploration--exploitation is a central trade-off in bandit learning. While classical algorithms such as upper confidence bound methods and Thompson Sampling effectively balance this trade-off in stationary environments, their exploration strategies mainly reduce uncertainty about the current optimal arm, which can be insufficient in nonstationary settings where future optimal arms may differ substantially from current ones. In this paper, we propose Future Information-Directed Sampling (FIDS), a new algorithm for Bayesian nonstationary bandits that explicitly explores to gather information about future optimal arms. We show that FIDS achieves regret comparable to Thompson Sampling up to a small constant factor, while being able to exploit predictive information structures that conventional exploration objectives fail to capture. To address the practical difficulty of posterior inference, we further propose a supervised-learning-based approximation framework that learns the FIDS policy from offline data, and demonstrate its effectiveness on synthetic benchmarks.