π€ AI Summary
This paper studies adversarial multi-armed bandit optimization for nonconvex, nonsmooth functions, where the loss in each round comprises a linear term plus an arbitrary small perturbation adaptively chosen after observing the playerβs action. We establish the first unified theoretical framework for adversarial bandits under approximately linear structure, integrating online gradient estimation, random directional sampling, adaptive perturbation analysis, and information-theoretic lower bound construction. Our analysis yields tight regret bounds: $O(sqrt{T})$ expected regret and $O(sqrt{T log T})$ high-probability regret, both matched by a $Omega(sqrt{T})$ lower bound. These results significantly improve upon prior high-probability analyses for linear bandits and provide novel upper bounds and theoretical guarantees for nonsmooth adversarial optimization.
π Abstract
We consider a bandit optimization problem for nonconvex and non-smooth functions, where in each trial the loss function is the sum of a linear function and a small but arbitrary perturbation chosen after observing the player's choice. We give both expected and high probability regret bounds for the problem. Our result also implies an improved high-probability regret bound for the bandit linear optimization, a special case with no perturbation. We also give a lower bound on the expected regret.