🤖 AI Summary
This paper studies the combinatorial semi-bandit problem with graph feedback under adversarial environments: in each round, the learner selects a subset of arms and observes rewards of both selected arms and their neighbors in a given feedback graph. For this novel setting, we establish the first tight regret bounds—both lower and upper—of order $widetilde{Theta}(Ssqrt{T} + sqrt{alpha S T})$, where $S$ is the action set size and $alpha$ is the independence number of the feedback graph, revealing their coupled impact on learning difficulty. We propose a convex relaxation framework based on negatively correlated randomization to effectively embed the discrete combinatorial action space into a continuous domain. Our theoretical analysis unifies full-information and standard semi-bandit settings as special cases. Furthermore, we provide constructive algorithms that achieve the derived bounds, thereby confirming their tightness and attainability.
📝 Abstract
In combinatorial semi-bandits, a learner repeatedly selects from a combinatorial decision set of arms, receives the realized sum of rewards, and observes the rewards of the individual selected arms as feedback. In this paper, we extend this framework to include emph{graph feedback}, where the learner observes the rewards of all neighboring arms of the selected arms in a feedback graph $G$. We establish that the optimal regret over a time horizon $T$ scales as $widetilde{Theta}(Ssqrt{T}+sqrt{alpha ST})$, where $S$ is the size of the combinatorial decisions and $alpha$ is the independence number of $G$. This result interpolates between the known regrets $widetildeTheta(Ssqrt{T})$ under full information (i.e., $G$ is complete) and $widetildeTheta(sqrt{KST})$ under the semi-bandit feedback (i.e., $G$ has only self-loops), where $K$ is the total number of arms. A key technical ingredient is to realize a convexified action using a random decision vector with negative correlations.