Adversarial bandit optimization for approximately linear functions

πŸ“… 2025-05-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF

career value

251K/year
πŸ€– AI Summary
This paper studies adversarial multi-armed bandit optimization for nonconvex, nonsmooth functions, where the loss in each round comprises a linear term plus an arbitrary small perturbation adaptively chosen after observing the player’s action. We establish the first unified theoretical framework for adversarial bandits under approximately linear structure, integrating online gradient estimation, random directional sampling, adaptive perturbation analysis, and information-theoretic lower bound construction. Our analysis yields tight regret bounds: $O(sqrt{T})$ expected regret and $O(sqrt{T log T})$ high-probability regret, both matched by a $Omega(sqrt{T})$ lower bound. These results significantly improve upon prior high-probability analyses for linear bandits and provide novel upper bounds and theoretical guarantees for nonsmooth adversarial optimization.

Technology Category

Application Category

πŸ“ Abstract
We consider a bandit optimization problem for nonconvex and non-smooth functions, where in each trial the loss function is the sum of a linear function and a small but arbitrary perturbation chosen after observing the player's choice. We give both expected and high probability regret bounds for the problem. Our result also implies an improved high-probability regret bound for the bandit linear optimization, a special case with no perturbation. We also give a lower bound on the expected regret.
Problem

Research questions and friction points this paper is trying to address.

Optimizing nonconvex nonsmooth functions with adversarial bandits
Bounding regret for linear functions with arbitrary perturbations
Improving high-probability regret bounds for bandit linear optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial bandit optimization for nonconvex functions
Regret bounds for linear and perturbed losses
Improved high-probability regret bounds
πŸ”Ž Similar Papers
No similar papers found.