Bayesian Analysis of Combinatorial Gaussian Process Bandits

📅 2023-12-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper studies the combinatorial, time-varying Gaussian process (GP) semi-bandit problem: at each round, a subset of dynamically available base arms is selected to maximize Bayesian cumulative reward. We first derive a tight Bayesian cumulative regret upper bound for GP-BayesUCB. We then unify and extend the regret analyses of GP-UCB and GP-Thompson sampling (TS) to infinite domains, time-varying structures, and combinatorial action spaces—while supporting contextual modeling. Theoretically, all three algorithms achieve $ ilde{O}(sqrt{Tgamma_T})$ regret, where $gamma_T$ denotes the maximum GP information gain up to horizon $T$. Empirically, on a real-world online energy-efficient navigation task, our methods significantly outperform existing baselines, validating both theoretical guarantees and practical utility.
📝 Abstract
We consider the combinatorial volatile Gaussian process (GP) semi-bandit problem. Each round, an agent is provided a set of available base arms and must select a subset of them to maximize the long-term cumulative reward. We study the Bayesian setting and provide novel Bayesian cumulative regret bounds for three GP-based algorithms: GP-UCB, GP-BayesUCB and GP-TS. Our bounds extend previous results for GP-UCB and GP-TS to the infinite, volatile and combinatorial setting, and to the best of our knowledge, we provide the first regret bound for GP-BayesUCB. Volatile arms encompass other widely considered bandit problems such as contextual bandits. Furthermore, we employ our framework to address the challenging real-world problem of online energy-efficient navigation, where we demonstrate its effectiveness compared to the alternatives.
Problem

Research questions and friction points this paper is trying to address.

Maximize long-term reward in combinatorial GP semi-bandits.
Extend Bayesian regret bounds for GP-based algorithms.
Apply framework to online energy-efficient navigation challenges.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bayesian analysis for GP bandits
Extending regret bounds
Online energy-efficient navigation framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Chalmers University of Technology | University of Gothenburg | Volvo Car Corporation
Jack Sandberg
Jack Sandberg
PhD student, Chalmers University of Technology
Multi-armed banditsReinforcment LearningMachine learning
N
Niklas Åkerblom
Volvo Car Corporation, Chalmers University of Technology and University of Gothenburg, Gothenburg, Sweden
M
M. Chehreghani
Chalmers University of Technology and University of Gothenburg, Gothenburg, Sweden