combinatorial semi-bandit learning

Designs and analyzes online learning algorithms that choose combinatorial actions (sets or combinations of atomic arms) under semi‑bandit feedback, including contextual variants that incorporate side information. Builds algorithms and formal regret analyses that operate with partial per‑component feedback and provide low‑regret guarantees — e.g., logarithmic pseudo‑regret in stochastic regimes and sublinear (e.g., o(√t)) regret in adversarial regimes.

combinatorialsemi-banditlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Adversarial Combinatorial Semi-bandits with Graph Feedback

Feb 26, 2025
YW
Yuxiao Wen
🏛️ New York University

This paper studies the combinatorial semi-bandit problem with graph feedback under adversarial environments: in each round, the learner selects a subset of arms and observes rewards of both selected arms and their neighbors in a given feedback graph. For this novel setting, we establish the first tight regret bounds—both lower and upper—of order $widetilde{Theta}(Ssqrt{T} + sqrt{alpha S T})$, where $S$ is the action set size and $alpha$ is the independence number of the feedback graph, revealing their coupled impact on learning difficulty. We propose a convex relaxation framework based on negatively correlated randomization to effectively embed the discrete combinatorial action space into a continuous domain. Our theoretical analysis unifies full-information and standard semi-bandit settings as special cases. Furthermore, we provide constructive algorithms that achieve the derived bounds, thereby confirming their tightness and attainability.

Determines optimal regret scaling with graph properties.Extends combinatorial semi-bandits with graph feedback.Interpolates between full information and semi-bandit feedback.

Oracle-Efficient Combinatorial Semi-Bandits

Oct 24, 2025
JK
Jung-hun Kim
🏛️ CREST | ENSAE | IP Paris | FairPlay joint team | London School of Economics | Seoul National University

This paper studies combinatorial semi-bandits, where an agent selects a subset of base arms per round and observes feedback from each selected arm. While practically important, existing algorithms rely on one expensive combinatorial optimization oracle call per round, severely limiting scalability. To address this, we propose a novel online learning framework that reduces the per-round oracle calls to only $O(log log T)$ while achieving the optimal $O(sqrt{T})$ regret bound. Our key contributions are: (1) a covariance-adaptive UCB strategy that explicitly models the reward noise structure; (2) a unified treatment accommodating both linear and nonlinear reward functions; and (3) tight theoretical guarantees under both worst-case and general smooth reward settings. Experiments demonstrate significant improvements in both computational efficiency and empirical performance.

Achieving sublinear regret with logarithmic oracle queriesExtending framework to nonlinear rewards with guaranteesReducing oracle calls in combinatorial semi-bandits

Multi-Play Combinatorial Semi-Bandit Problem

Sep 11, 2025
SN
Shintaro Nakamura
🏛️ University of Tokyo | CENTAI Institute | Microsoft Research

Existing combinatorial semi-bandits (CSBs) are restricted to binary actions, limiting their applicability to fundamental combinatorial optimization problems such as optimal transport and knapsack, which require nonnegative integer-valued action vectors. Method: We propose the Multi-Choice Combinatorial Semi-Bandit (MP-CSB) framework—the first to generalize action spaces to nonnegative integer vectors—and design an efficient Thompson sampling–based algorithm. To ensure robustness against both stochastic and adversarial environments, we introduce a “dual-robust” algorithm integrating variance-adaptive analysis, path-length control, and quadratic variation techniques to handle exponential action spaces and heterogeneous feedback. Contribution/Results: We establish tight theoretical guarantees: an $O(log T)$ distribution-dependent regret under stochastic rewards and a $ ilde{mathcal{O}}(sqrt{T})$ worst-case regret under adversarial rewards. Empirical evaluation demonstrates significant improvements over state-of-the-art CSB methods across diverse combinatorial optimization benchmarks.

Develops algorithms for stochastic and adversarial regret regimesExtends combinatorial bandits to non-negative integer action spacesSolves limitations in optimal transport and knapsack problems

Near-Optimal Regret for Efficient Stochastic Combinatorial Semi-Bandits

Aug 08, 2025
ZY
Zichun Ye
🏛️ Shanghai Jiao Tong University | Carnegie Mellon University

In combinatorial multi-armed bandits (CMAB), UCB-type algorithms incur an undesirable $O(log T)$ regret overhead, while adversarial approaches (e.g., EXP3.M) suffer from excessive computational cost. Method: This paper proposes an efficient randomized combinatorial semi-bandit decision framework and the CMOSS algorithm, which innovatively integrates combinatorial optimization, stochastic modeling, and a hybrid semi-/cascade feedback mechanism, grounded in a minimax-optimal strategy design. Contribution/Results: CMOSS guarantees polynomial-time solvability while completely eliminating the $log T$ factor in regret. Its cumulative regret is theoretically bounded by $Oig((log k)^2 sqrt{kmT}ig)$, nearly matching the lower bound $Omega(sqrt{kmT})$. Extensive experiments on synthetic and real-world datasets demonstrate that CMOSS significantly outperforms baseline methods—including UCB variants and EXP3.M—in both regret performance and computational efficiency.

Achieves near-optimal regret with semi-bandit feedbackEliminates detrimental logarithmic regret factor dependencyResolves computational inefficiency in combinatorial bandits

Contextual Combinatorial Bandits with Changing Action Sets via Gaussian Processes

Oct 05, 2021
AN
Andi Nika
🏛️ Max Planck Institute for Software Systems | EPFL | Bilkent University

This paper studies the contextual combinatorial bandit problem with a dynamically evolving base arm set over time, aiming to maximize cumulative reward. To address the dual challenges of time-varying feasible action sets and context-dependent rewards, we introduce Gaussian process (GP) modeling into this framework for the first time, proposing the O’CLOK-UCB algorithm and its sparse GP-accelerated variant. Our method integrates kernelized UCB, combinatorial feasibility constraints, and Lipschitz continuity analysis. We establish a sublinear regret bound of Õ(√(λ∗(K)KTγ_T)), where λ∗(K) is the largest eigenvalue of the action covariance matrix and γ_T is the maximum information gain—revealing their coupled impact on regret. Empirical evaluation on real-world datasets demonstrates significant improvements over existing UCB-based approaches, confirming both theoretical rigor and practical efficacy.

Modeling base arm outcomes via Gaussian Process contextual dependenciesOptimizing cumulative reward under changing action set constraintsSolving combinatorial bandits with time-varying available actions

Latest Papers

What's happening recently
View more

This work addresses the contextual combinatorial semi-bandit (CCSB) problem, where at each round only contextual information is observed, and the learner must select a combinatorial action satisfying a cardinality constraint to maximize cumulative reward, without assuming any structural properties of the action space. To tackle this setting, the authors propose SquareCB.Comb, an algorithm that balances exploration and exploitation by solving a convex optimization problem at each round and efficiently samples combinatorial actions. This method achieves, for the first time under general function approximation and arbitrary combinatorial action structures, a minimax-optimal regret bound of $O(\sqrt{m A T \log|\mathcal{F}|})$, without requiring additional assumptions on the action set. In the realizable setting, it matches the performance of the best existing policy search approaches while offering superior generalization capabilities.

combinatorial actioncontextual combinatorial semi-banditscumulative reward maximization

This work addresses the adversarial multi-armed bandit problem under partial monitoring, where losses of unplayed actions are independently revealed with an unknown probability \( r \), corresponding to an Erdős–Rényi side-observation graph. The paper introduces the first adaptive algorithmic framework that achieves near-optimal regret bounds without prior knowledge of \( r \). Specifically, it proposes two algorithms: when \( r \geq \frac{\log T}{2N} \), the expected regret is \( O(\sqrt{(T/r)\log N}) \); for smaller \( r \), the regret bound becomes \( O(\sqrt{(T/r)\log(N+T)}) \). A fast estimation mechanism automatically identifies the regime of \( r \), enabling the framework to match the known-\( r \) lower bound up to logarithmic factors.

adversarial multi-armed banditsErdős-Rényi graphsonline learning

This study addresses the multi-armed bandit problem augmented with an oracle that can reveal the optimal action at a cost, investigating whether such query capability reduces regret under standard bandit feedback where only the reward of the chosen action is observed. By integrating information-theoretic lower bounds, stochastic process analysis, and adaptive algorithm design, the work provides the first complete characterization of the value of this querying mechanism and uncovers fundamental differences between adversarial or correlated environments and i.i.d. settings. The main contributions include establishing a regret lower bound of Ω(√(T−k)) in adversarial or correlated environments, and achieving matching upper and lower bounds of Õ(min{T/k, √(T−k)}) in the i.i.d. case, thereby rigorously quantifying how the number of queries k fundamentally governs learning performance.

bandit feedbackbest-action queriesmulti-armed bandits

This work addresses the exponential blow-up in action space inherent in adversarial combinatorial multi-armed bandits, where at each round an agent selects $m$ items out of $d$ and observes only an aggregated loss. The paper proposes an efficient algorithm that exploits the structural assumption that the loss is determined by a $d$-dimensional item loss vector, thereby avoiding explicit enumeration of all $\binom{d}{m}$ actions. By introducing a dual representation and parameterizing a low-dimensional sampling distribution, the method integrates online learning with combinatorial optimization to achieve, for the first time, a high-probability regret bound of $O(\sqrt{dT \log(K/\delta)})$ with probability at least $1-\delta$ (where $K = \binom{d}{m}$) in polynomial time. This matches the theoretical performance of EXP3-KW while eliminating exponential space complexity, resolving an open problem posed by Maiti et al.

adversarial banditscombinatorial banditslarge action space

This work addresses the challenge of leveraging tree-structured action similarities—encoded such that the loss function satisfies tree compatibility—in multi-armed bandit problems, where traditional single-point feedback fails to effectively exploit this structure to reduce regret. The authors propose a unified adaptive online learning algorithm that accommodates a spectrum of multi-point feedback mechanisms, ranging from semi-bandit to minimal two-point feedback. By introducing a similarity-aware effective number of actions, denoted $K_{\text{eff}}$, to replace the original action count $K$, the algorithm achieves improved regret bounds. Theoretical analysis reveals a fundamental limitation of single-point feedback in harnessing tree similarity and establishes, for the first time, an optimal $\sqrt{T}$ regret bound for Lipschitz bandits with dimension $d \leq 2$ under two-point feedback, striking an optimal balance between generality and structural exploitation.

bandit feedbackmulti-armed banditsonline learning

Hot Scholars

VA

Vaneet Aggarwal

Professor and University Faculty Scholar, Purdue University
Machine LearningReinforcement LearningQuantum ComputingNetworking
MH

Min-hwan Oh

Seoul National University
Reinforcement LearningBandit AlgorithmsMachine Learning
XW

Xuchuang Wang

UMass Amherst
Performance EvaluationOnline OptimizationQuantum NetworkMulti-Agent System
XL

Xutong Liu

Assistant Professor of Computer Science and Systems, University of Washington
Reinforcement LearningOnline LearningCombinatorial OptimizationNetwork Systems
JS

Jack Sandberg

PhD student, Chalmers University of Technology
Multi-armed banditsReinforcment LearningMachine learning