๐ค AI Summary
This paper studies the Lipschitz bandit problem under stochastic delayed feedback: the action space is a metric space with Lipschitz-continuous expected rewards, and reward observations incur i.i.d. random delaysโeither bounded or unbounded. To address this novel setting, we propose a delay-aware zooming algorithm for bounded delays and a phased learning strategy for unbounded delays, achieving the first sublinear regret bounds of $O(T^{(d+1)/(d+2)}log T)$ and $O(T^{(d+2)/(d+3)}log T)$, respectively, where $d$ denotes the metric space dimension. Our theoretical analysis establishes matching lower bounds, confirming near-optimality. Experiments demonstrate robustness and efficiency across diverse delay distributions. The core innovation lies in tightly coupling Lipschitz structure with dynamic modeling of delay effects, enabling principled continuous decision-making under information lag.
๐ Abstract
The Lipschitz bandit problem extends stochastic bandits to a continuous action set defined over a metric space, where the expected reward function satisfies a Lipschitz condition. In this work, we introduce a new problem of Lipschitz bandit in the presence of stochastic delayed feedback, where the rewards are not observed immediately but after a random delay. We consider both bounded and unbounded stochastic delays, and design algorithms that attain sublinear regret guarantees in each setting. For bounded delays, we propose a delay-aware zooming algorithm that retains the optimal performance of the delay-free setting up to an additional term that scales with the maximal delay $ฯ_{max}$. For unbounded delays, we propose a novel phased learning strategy that accumulates reliable feedback over carefully scheduled intervals, and establish a regret lower bound showing that our method is nearly optimal up to logarithmic factors. Finally, we present experimental results to demonstrate the efficiency of our algorithms under various delay scenarios.