Adversarial Bandit Optimization with Globally Bounded Perturbations to Linear Losses

📅 2026-03-27
📈 Citations: 0
Influential: 0
📄 PDF

career value

238K/year
🤖 AI Summary
This work addresses adversarial bandit optimization under a global perturbation budget, where in each round the loss consists of a linear function plus an action-dependent perturbation term, with the total perturbation constrained globally. Within this non-convex and non-smooth setting, the paper establishes—for the first time—both expected and high-probability regret upper bounds under a global perturbation budget, improving upon the classical high-probability regret bound in the unperturbed case. Additionally, it provides a matching lower bound on the expected regret. The analysis combines techniques from adversarial bandits, perturbation modeling, and refined probabilistic arguments, offering rigorous theoretical guarantees for online decision-making in perturbed environments.

Technology Category

Application Category

📝 Abstract
We study a class of adversarial bandit optimization problems in which the loss functions may be non-convex and non-smooth. In each round, the learner observes a loss that consists of an underlying linear component together with an additional perturbation applied after the learner selects an action. The perturbations are measured relative to the linear losses and are constrained by a global budget that bounds their cumulative magnitude over time. Under this model, we establish both expected and high-probability regret guarantees. As a special case of our analysis, we recover an improved high-probability regret bound for classical bandit linear optimization, which corresponds to the setting without perturbations. We further complement our upper bounds by proving a lower bound on the expected regret.
Problem

Research questions and friction points this paper is trying to address.

adversarial bandit
non-convex loss
non-smooth loss
globally bounded perturbations
linear losses
Innovation

Methods, ideas, or system contributions that make the work stand out.

adversarial bandits
globally bounded perturbations
linear losses
regret bounds
non-convex optimization
🔎 Similar Papers
No similar papers found.
Z
Zhuoyu Cheng
Joint Graduate School of Mathematics for Innovation, Kyushu University, Japan.
K
Kohei Hatano
Department of Informatics, Kyushu University, Japan. and RIKEN AIP, Japan.
E
Eiji Takimoto
Department of Informatics, Kyushu University, Japan.