Efficient Prior Selection in Gaussian Process Bandits with Thompson Sampling

๐Ÿ“… 2025-02-03
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This paper addresses the challenge of joint prior learning and optimization in Gaussian process (GP) bandits when the true prior is unknown, aiming to minimize cumulative regret induced by prior misspecification. To overcome the lack of theoretical guarantees in conventional maximum likelihood estimation (MLE)-based prior selection, we propose two novel GP-Thompson samplingโ€“based algorithms: Prior-Elimination, which employs online prior selection via confidence-set elimination, and HyperPrior, which adaptively updates the prior through Bayesian hierarchical modeling with a hyperprior. We establish the first joint prior learning framework for GP bandits with provably sublinear regret bounds, rigorously proving $ ilde{O}(sqrt{Tgamma_T})$ regret for both algorithms under general kernels, where $gamma_T$ denotes the maximum information gain. Extensive experiments on synthetic and real-world datasets demonstrate significant improvements over existing baselines.

Technology Category

Machine Learning: Online Learning & BanditsReasoning under Uncertainty: Stochastic OptimizationSearch and Optimization: Learning to Search

Application Category

Graph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
๐Ÿ“ Abstract
Gaussian process (GP) bandits provide a powerful framework for solving blackbox optimization of unknown functions. The characteristics of the unknown function depends heavily on the assumed GP prior. Most work in the literature assume that this prior is known but in practice this seldom holds. Instead, practitioners often rely on maximum likelihood estimation to select the hyperparameters of the prior - which lacks theoretical guarantees. In this work, we propose two algorithms for joint prior selection and regret minimization in GP bandits based on GP Thompson sampling (GP-TS): Prior-Elimination GP-TS (PE-GP-TS) and HyperPrior GP-TS (HP-GP-TS). We theoretically analyze the algorithms and establish upper bounds for their respective regret. In addition, we demonstrate the effectiveness of our algorithms compared to the alternatives through experiments with synthetic and real-world data.
Problem

Research questions and friction points this paper is trying to address.

Thompson Sampling
Gaussian Processes
Prior Selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gaussian Process Bandits
Prior Optimization
Thompson Sampling
๐Ÿ”Ž Similar Papers
Jack Sandberg
Jack Sandberg
PhD student, Chalmers University of Technology
Multi-armed banditsReinforcment LearningMachine learning
M
M. Chehreghani
Chalmers University of Technology and University of Gothenburg, Gothenburg, Sweden