Lipschitz Bandits with Stochastic Delayed Feedback

๐Ÿ“… 2025-09-30
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This paper studies the Lipschitz bandit problem under stochastic delayed feedback: the action space is a metric space with Lipschitz-continuous expected rewards, and reward observations incur i.i.d. random delaysโ€”either bounded or unbounded. To address this novel setting, we propose a delay-aware zooming algorithm for bounded delays and a phased learning strategy for unbounded delays, achieving the first sublinear regret bounds of $O(T^{(d+1)/(d+2)}log T)$ and $O(T^{(d+2)/(d+3)}log T)$, respectively, where $d$ denotes the metric space dimension. Our theoretical analysis establishes matching lower bounds, confirming near-optimality. Experiments demonstrate robustness and efficiency across diverse delay distributions. The core innovation lies in tightly coupling Lipschitz structure with dynamic modeling of delay effects, enabling principled continuous decision-making under information lag.

Technology Category

Machine Learning: Online Learning & BanditsSearch and Optimization: Learning to SearchReasoning under Uncertainty: Stochastic Optimization

Application Category

Graph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsResponsible Web: Human-perceived consequences of algorithmic deployment on the webEconomics, Online Markets and Human Computation: Uses of LLMs and GenAI for marketplace design, bidding, and strategic interactions
๐Ÿ“ Abstract
The Lipschitz bandit problem extends stochastic bandits to a continuous action set defined over a metric space, where the expected reward function satisfies a Lipschitz condition. In this work, we introduce a new problem of Lipschitz bandit in the presence of stochastic delayed feedback, where the rewards are not observed immediately but after a random delay. We consider both bounded and unbounded stochastic delays, and design algorithms that attain sublinear regret guarantees in each setting. For bounded delays, we propose a delay-aware zooming algorithm that retains the optimal performance of the delay-free setting up to an additional term that scales with the maximal delay $ฯ„_{max}$. For unbounded delays, we propose a novel phased learning strategy that accumulates reliable feedback over carefully scheduled intervals, and establish a regret lower bound showing that our method is nearly optimal up to logarithmic factors. Finally, we present experimental results to demonstrate the efficiency of our algorithms under various delay scenarios.
Problem

Research questions and friction points this paper is trying to address.

Lipschitz bandits with stochastic delayed reward feedback
Algorithms for bounded and unbounded stochastic delay settings
Achieving sublinear regret guarantees under delayed observations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Delay-aware zooming algorithm for bounded delays
Phased learning strategy for unbounded delays
Sublinear regret guarantees under stochastic delays
๐Ÿ”Ž Similar Papers
Z
Zhongxuan Liu
Department of Statistics, University of California, Davis, Davis, CA 95616
Y
Yue Kang
Department of Statistics, University of California, Davis, Davis, CA 95616
T
Thomas C. M. Lee
Department of Statistics, University of California, Davis, Davis, CA 95616