Adversarial Combinatorial Semi-bandits with Graph Feedback

📅 2025-02-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper studies the combinatorial semi-bandit problem with graph feedback under adversarial environments: in each round, the learner selects a subset of arms and observes rewards of both selected arms and their neighbors in a given feedback graph. For this novel setting, we establish the first tight regret bounds—both lower and upper—of order $widetilde{Theta}(Ssqrt{T} + sqrt{alpha S T})$, where $S$ is the action set size and $alpha$ is the independence number of the feedback graph, revealing their coupled impact on learning difficulty. We propose a convex relaxation framework based on negatively correlated randomization to effectively embed the discrete combinatorial action space into a continuous domain. Our theoretical analysis unifies full-information and standard semi-bandit settings as special cases. Furthermore, we provide constructive algorithms that achieve the derived bounds, thereby confirming their tightness and attainability.

Technology Category

Machine Learning: Online Learning & BanditsGame Theory and Economic Paradigms: Adversarial LearningSearch and Optimization: Combinatorial Optimization

Application Category

Graph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: Social networks and social learning
📝 Abstract
In combinatorial semi-bandits, a learner repeatedly selects from a combinatorial decision set of arms, receives the realized sum of rewards, and observes the rewards of the individual selected arms as feedback. In this paper, we extend this framework to include emph{graph feedback}, where the learner observes the rewards of all neighboring arms of the selected arms in a feedback graph $G$. We establish that the optimal regret over a time horizon $T$ scales as $widetilde{Theta}(Ssqrt{T}+sqrt{alpha ST})$, where $S$ is the size of the combinatorial decisions and $alpha$ is the independence number of $G$. This result interpolates between the known regrets $widetildeTheta(Ssqrt{T})$ under full information (i.e., $G$ is complete) and $widetildeTheta(sqrt{KST})$ under the semi-bandit feedback (i.e., $G$ has only self-loops), where $K$ is the total number of arms. A key technical ingredient is to realize a convexified action using a random decision vector with negative correlations.
Problem

Research questions and friction points this paper is trying to address.

Extends combinatorial semi-bandits with graph feedback.
Determines optimal regret scaling with graph properties.
Interpolates between full information and semi-bandit feedback.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial combinatorial semi-bandits
Graph feedback mechanism
Convexified action realization
🔎 Similar Papers
2024-07-24arXiv.orgCitations: 4
💼 Related Jobs
No related jobs found.
Y
Yuxiao Wen
Courant Institute of Mathematical Sciences, New York University