๐ค AI Summary
This paper addresses the challenge of joint prior learning and optimization in Gaussian process (GP) bandits when the true prior is unknown, aiming to minimize cumulative regret induced by prior misspecification. To overcome the lack of theoretical guarantees in conventional maximum likelihood estimation (MLE)-based prior selection, we propose two novel GP-Thompson samplingโbased algorithms: Prior-Elimination, which employs online prior selection via confidence-set elimination, and HyperPrior, which adaptively updates the prior through Bayesian hierarchical modeling with a hyperprior. We establish the first joint prior learning framework for GP bandits with provably sublinear regret bounds, rigorously proving $ ilde{O}(sqrt{Tgamma_T})$ regret for both algorithms under general kernels, where $gamma_T$ denotes the maximum information gain. Extensive experiments on synthetic and real-world datasets demonstrate significant improvements over existing baselines.
๐ Abstract
Gaussian process (GP) bandits provide a powerful framework for solving blackbox optimization of unknown functions. The characteristics of the unknown function depends heavily on the assumed GP prior. Most work in the literature assume that this prior is known but in practice this seldom holds. Instead, practitioners often rely on maximum likelihood estimation to select the hyperparameters of the prior - which lacks theoretical guarantees. In this work, we propose two algorithms for joint prior selection and regret minimization in GP bandits based on GP Thompson sampling (GP-TS): Prior-Elimination GP-TS (PE-GP-TS) and HyperPrior GP-TS (HP-GP-TS). We theoretically analyze the algorithms and establish upper bounds for their respective regret. In addition, we demonstrate the effectiveness of our algorithms compared to the alternatives through experiments with synthetic and real-world data.