Score
Designs and analyzes online learning algorithms that choose combinatorial actions (sets or combinations of atomic arms) under semi‑bandit feedback, including contextual variants that incorporate side information. Builds algorithms and formal regret analyses that operate with partial per‑component feedback and provide low‑regret guarantees — e.g., logarithmic pseudo‑regret in stochastic regimes and sublinear (e.g., o(√t)) regret in adversarial regimes.
This paper studies the combinatorial semi-bandit problem with graph feedback under adversarial environments: in each round, the learner selects a subset of arms and observes rewards of both selected arms and their neighbors in a given feedback graph. For this novel setting, we establish the first tight regret bounds—both lower and upper—of order $widetilde{Theta}(Ssqrt{T} + sqrt{alpha S T})$, where $S$ is the action set size and $alpha$ is the independence number of the feedback graph, revealing their coupled impact on learning difficulty. We propose a convex relaxation framework based on negatively correlated randomization to effectively embed the discrete combinatorial action space into a continuous domain. Our theoretical analysis unifies full-information and standard semi-bandit settings as special cases. Furthermore, we provide constructive algorithms that achieve the derived bounds, thereby confirming their tightness and attainability.
This paper studies combinatorial semi-bandits, where an agent selects a subset of base arms per round and observes feedback from each selected arm. While practically important, existing algorithms rely on one expensive combinatorial optimization oracle call per round, severely limiting scalability. To address this, we propose a novel online learning framework that reduces the per-round oracle calls to only $O(log log T)$ while achieving the optimal $O(sqrt{T})$ regret bound. Our key contributions are: (1) a covariance-adaptive UCB strategy that explicitly models the reward noise structure; (2) a unified treatment accommodating both linear and nonlinear reward functions; and (3) tight theoretical guarantees under both worst-case and general smooth reward settings. Experiments demonstrate significant improvements in both computational efficiency and empirical performance.
Existing combinatorial semi-bandits (CSBs) are restricted to binary actions, limiting their applicability to fundamental combinatorial optimization problems such as optimal transport and knapsack, which require nonnegative integer-valued action vectors. Method: We propose the Multi-Choice Combinatorial Semi-Bandit (MP-CSB) framework—the first to generalize action spaces to nonnegative integer vectors—and design an efficient Thompson sampling–based algorithm. To ensure robustness against both stochastic and adversarial environments, we introduce a “dual-robust” algorithm integrating variance-adaptive analysis, path-length control, and quadratic variation techniques to handle exponential action spaces and heterogeneous feedback. Contribution/Results: We establish tight theoretical guarantees: an $O(log T)$ distribution-dependent regret under stochastic rewards and a $ ilde{mathcal{O}}(sqrt{T})$ worst-case regret under adversarial rewards. Empirical evaluation demonstrates significant improvements over state-of-the-art CSB methods across diverse combinatorial optimization benchmarks.
In combinatorial multi-armed bandits (CMAB), UCB-type algorithms incur an undesirable $O(log T)$ regret overhead, while adversarial approaches (e.g., EXP3.M) suffer from excessive computational cost. Method: This paper proposes an efficient randomized combinatorial semi-bandit decision framework and the CMOSS algorithm, which innovatively integrates combinatorial optimization, stochastic modeling, and a hybrid semi-/cascade feedback mechanism, grounded in a minimax-optimal strategy design. Contribution/Results: CMOSS guarantees polynomial-time solvability while completely eliminating the $log T$ factor in regret. Its cumulative regret is theoretically bounded by $Oig((log k)^2 sqrt{kmT}ig)$, nearly matching the lower bound $Omega(sqrt{kmT})$. Extensive experiments on synthetic and real-world datasets demonstrate that CMOSS significantly outperforms baseline methods—including UCB variants and EXP3.M—in both regret performance and computational efficiency.
This paper studies the contextual combinatorial bandit problem with a dynamically evolving base arm set over time, aiming to maximize cumulative reward. To address the dual challenges of time-varying feasible action sets and context-dependent rewards, we introduce Gaussian process (GP) modeling into this framework for the first time, proposing the O’CLOK-UCB algorithm and its sparse GP-accelerated variant. Our method integrates kernelized UCB, combinatorial feasibility constraints, and Lipschitz continuity analysis. We establish a sublinear regret bound of Õ(√(λ∗(K)KTγ_T)), where λ∗(K) is the largest eigenvalue of the action covariance matrix and γ_T is the maximum information gain—revealing their coupled impact on regret. Empirical evaluation on real-world datasets demonstrates significant improvements over existing UCB-based approaches, confirming both theoretical rigor and practical efficacy.
This work addresses the contextual combinatorial semi-bandit (CCSB) problem, where at each round only contextual information is observed, and the learner must select a combinatorial action satisfying a cardinality constraint to maximize cumulative reward, without assuming any structural properties of the action space. To tackle this setting, the authors propose SquareCB.Comb, an algorithm that balances exploration and exploitation by solving a convex optimization problem at each round and efficiently samples combinatorial actions. This method achieves, for the first time under general function approximation and arbitrary combinatorial action structures, a minimax-optimal regret bound of $O(\sqrt{m A T \log|\mathcal{F}|})$, without requiring additional assumptions on the action set. In the realizable setting, it matches the performance of the best existing policy search approaches while offering superior generalization capabilities.
This work addresses the adversarial multi-armed bandit problem under partial monitoring, where losses of unplayed actions are independently revealed with an unknown probability \( r \), corresponding to an Erdős–Rényi side-observation graph. The paper introduces the first adaptive algorithmic framework that achieves near-optimal regret bounds without prior knowledge of \( r \). Specifically, it proposes two algorithms: when \( r \geq \frac{\log T}{2N} \), the expected regret is \( O(\sqrt{(T/r)\log N}) \); for smaller \( r \), the regret bound becomes \( O(\sqrt{(T/r)\log(N+T)}) \). A fast estimation mechanism automatically identifies the regime of \( r \), enabling the framework to match the known-\( r \) lower bound up to logarithmic factors.
This study addresses the multi-armed bandit problem augmented with an oracle that can reveal the optimal action at a cost, investigating whether such query capability reduces regret under standard bandit feedback where only the reward of the chosen action is observed. By integrating information-theoretic lower bounds, stochastic process analysis, and adaptive algorithm design, the work provides the first complete characterization of the value of this querying mechanism and uncovers fundamental differences between adversarial or correlated environments and i.i.d. settings. The main contributions include establishing a regret lower bound of Ω(√(T−k)) in adversarial or correlated environments, and achieving matching upper and lower bounds of Õ(min{T/k, √(T−k)}) in the i.i.d. case, thereby rigorously quantifying how the number of queries k fundamentally governs learning performance.
This work addresses the exponential blow-up in action space inherent in adversarial combinatorial multi-armed bandits, where at each round an agent selects $m$ items out of $d$ and observes only an aggregated loss. The paper proposes an efficient algorithm that exploits the structural assumption that the loss is determined by a $d$-dimensional item loss vector, thereby avoiding explicit enumeration of all $\binom{d}{m}$ actions. By introducing a dual representation and parameterizing a low-dimensional sampling distribution, the method integrates online learning with combinatorial optimization to achieve, for the first time, a high-probability regret bound of $O(\sqrt{dT \log(K/\delta)})$ with probability at least $1-\delta$ (where $K = \binom{d}{m}$) in polynomial time. This matches the theoretical performance of EXP3-KW while eliminating exponential space complexity, resolving an open problem posed by Maiti et al.
This work addresses the challenge of leveraging tree-structured action similarities—encoded such that the loss function satisfies tree compatibility—in multi-armed bandit problems, where traditional single-point feedback fails to effectively exploit this structure to reduce regret. The authors propose a unified adaptive online learning algorithm that accommodates a spectrum of multi-point feedback mechanisms, ranging from semi-bandit to minimal two-point feedback. By introducing a similarity-aware effective number of actions, denoted $K_{\text{eff}}$, to replace the original action count $K$, the algorithm achieves improved regret bounds. Theoretical analysis reveals a fundamental limitation of single-point feedback in harnessing tree similarity and establishes, for the first time, an optimal $\sqrt{T}$ regret bound for Lipschitz bandits with dimension $d \leq 2$ under two-point feedback, striking an optimal balance between generality and structural exploitation.