m-Set Adversarial Bandits with Winner Feedback

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of establishing regret bounds for m-set adversarial bandits under diverse utility and feedback models. To overcome the theoretical limitations of conventional combinatorial and multinomial logit (MNL) bandits, this work proposes an information-theoretic framework for analyzing regret lower bounds. It systematically derives matching upper and lower bounds across various feedback scenarios, with the theoretical findings validated through synthetic data experiments. The primary contributions include revealing how subtle variations in problem settings significantly impact learning rates, establishing rigorous theoretical boundaries, and empirically confirming the validity of the analysis. Ultimately, this research deepens the fundamental understanding of feedback mechanisms in adversarial bandit problems.
📝 Abstract
We show upper and lower bounds on the regret of $m$-set adversarial bandits for different utilities (winner reward or sum of rewards) and feedback models (winner index, winner reward, sum of rewards, and their combinations). By comparing to standard bounds for combinatorial and MNL bandits, our results reveal how subtle changes in the setting can have a dramatic impact on the learning rates. Our main technical contributions are the information-theoretic lower bounds on the regret. Experiments on synthetic data confirm our theoretical analyses.
Problem

Research questions and friction points this paper is trying to address.

m-set adversarial bandits
regret bounds
winner feedback
combinatorial bandits
online learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

m-set adversarial bandits
regret bounds
information-theoretic lower bounds
feedback models
learning rates
🔎 Similar Papers
No similar papers found.