Beyond Reward Suppression: Near-Optimal Offline Attacks on Warm-Start Bandits with Bounded Rewards

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing offline attacks on multi-armed bandits that solely suppress non-target arms, investigating target-promotion attacks in warm-start scenarios under bounded rewards. Theoretically, we prove the necessity of directly promoting the target arm under specific conditions, thereby transcending the conventional pure-suppression paradigm. Methodologically, we propose an optimal allocation strategy that balances target promotion with non-target suppression, and design a near-optimal, low-cost data injection algorithm. This work achieves sublinear attack costs while successfully misleading mainstream algorithms, including UCB and Thompson Sampling, into selecting the target arm. Extensive experiments on both real-world and synthetic datasets validate the effectiveness of the proposed attack framework.
📝 Abstract
Adversarial attacks on bandits aim to mislead a learner toward a target arm while keeping the attack cost small. Existing attacks typically achieve this by suppressing non-target arms. In practice, however, manipulation such as fake reviews often directly promotes the target item. We study this gap through bounded offline attacks on warm-start bandits, where an attacker can inject only valid action-reward pairs into the warm-start history before deployment. We show that target promotion is not merely a heuristic: when the target arm lies near the lower reward boundary, any order-optimal-cost attack against UCB that makes it selected in nearly all online rounds must allocate a nonvanishing fraction of its cost to the target arm. We then design an attack that achieves the optimal sublinear cost and characterize its allocation between target promotion and non-target suppression. We further extend the attack to Thompson Sampling, $ε$-greedy, and a broader class of bandit algorithms. Experiments on real-world and synthetic data validate the effectiveness of our attacks.
Problem

Research questions and friction points this paper is trying to address.

adversarial attacks
warm-start bandits
offline attacks
target promotion
bounded rewards
Innovation

Methods, ideas, or system contributions that make the work stand out.

Offline Adversarial Attacks
Warm-Start Bandits
Target Promotion
Bounded Rewards
Sublinear Attack Cost