🤖 AI Summary
This paper studies the non-stationary $K$-armed bandit problem with Hölder-continuous reward functions—characterized by smoothness exponent $eta$ and coefficient $lambda$—and time-varying rewards, focusing on dynamic regret minimization. We propose the first fully adaptive algorithm that requires no prior knowledge of $eta$ or $lambda$, integrating sliding-window estimation, adaptive segmentation, and a novel “safe-arm” mechanism. Our theoretical contributions are threefold: (1) We establish a tight minimax lower bound $Thetaig(T^{(1+eta)/(1+2eta)} lambda^{1/(1+2eta)}ig)$ for dynamic regret in general Hölder non-stationary environments; (2) We achieve full-range adaptivity—optimal for all $eta > 0$—for the first time; (3) We show that the safe-arm mechanism enables gap-dependent dynamic regret $O(log T)$, breaking the conventional $Omega(sqrt{T})$ barrier. These results resolve the long-standing open problem of adaptive dynamic regret control.
📝 Abstract
We study a $K$-armed non-stationary bandit model where rewards change smoothly, as captured by H""{o}lder class assumptions on rewards as functions of time. Such smooth changes are parametrized by a H""{o}lder exponent $eta$ and coefficient $lambda$. While various sub-cases of this general model have been studied in isolation, we first establish the minimax dynamic regret rate generally for all $K,eta,lambda$. Next, we show this optimal dynamic regret can be attained adaptively, without knowledge of $eta,lambda$. To contrast, even with parameter knowledge, upper bounds were only previously known for limited regimes $etaleq 1$ and $eta=2$ (Slivkins, 2014; Krishnamurthy and Gopalan, 2021; Manegueu et al., 2021; Jia et al.,2023). Thus, our work resolves open questions raised by these disparate threads of the literature. We also study the problem of attaining faster gap-dependent regret rates in non-stationary bandits. While such rates are long known to be impossible in general (Garivier and Moulines, 2011), we show that environments admitting a safe arm (Suk and Kpotufe, 2022) allow for much faster rates than the worst-case scaling with $sqrt{T}$. While previous works in this direction focused on attaining the usual logarithmic regret bounds, as summed over stationary periods, our new gap-dependent rates reveal new optimistic regimes of non-stationarity where even the logarithmic bounds are pessimistic. We show our new gap-dependent rate is tight and that its achievability (i.e., as made possible by a safe arm) has a surprisingly simple and clean characterization within the smooth H""{o}lder class model.