Score
Design, build, and evaluate online decision-making algorithms and policies that select actions given observed context to maximize cumulative reward, covering multi-armed, contextual, linear, and slate bandit formulations and their stochastic and delayed-feedback variants. Work includes constructing batched, rarely-switching, and limited-adaptivity procedures, formulating experimental-design-based exploration, implementing efficient per-round policy computation, and proving sublinear regret and other performance guarantees under streaming or delayed feedback.
This work addresses the lack of high-probability optimal regret bounds for policy optimization methods in stochastic contextual multi-armed bandits (CMAB). To bridge the gap between theoretical guarantees and practical applicability, the paper proposes an efficient algorithm that integrates policy optimization with general offline function approximation. The method establishes, for the first time, a high-probability optimal regret bound of $\widetilde{O}(\sqrt{K|\mathcal{A}| \log|\mathcal{F}|})$ for policy optimization in CMAB, where $K$ denotes the number of contexts, $|\mathcal{A}|$ the number of actions, and $|\mathcal{F}|$ the complexity of the policy class. Both theoretical analysis and empirical experiments corroborate the algorithm’s effectiveness and optimality, thereby providing a solid foundation for deploying policy optimization in real-world contextual bandit settings.
This paper addresses the problem of safely and efficiently improving online learning performance for stochastic multi-armed bandits (MAB) and combinatorial MAB under offline–online distribution shift. To tackle distribution mismatch between offline and online data, we propose MIN-UCB—a novel adaptive algorithm that automatically discards low-quality offline data when the shift magnitude is unknown, ensuring safety, and actively leverages bounded-shift offline information to substantially reduce regret when the shift is bounded. We establish, for the first time, a theoretical lower bound on regret for MAB with biased offline data. We prove that MIN-UCB achieves tight instance-dependent and instance-independent regret bounds—strictly outperforming classical UCB. Both theoretical analysis and numerical experiments demonstrate its robustness and superiority across diverse shift regimes.
This paper studies the nonparametric contextual bandit problem under batch constraints, where the reward function is an unknown smooth function of covariates and the policy is updated only upon completion of each batch. To address the exploration–exploitation trade-off inherent in batched learning, we propose a dynamic binning mechanism: bin widths adaptively scale with batch sizes, integrated with nonparametric regression and minimax analysis to achieve efficient estimation within the batched learning framework. We theoretically establish that only a constant number of policy updates suffice to attain the optimal online regret bound—up to logarithmic factors—and provide a matching lower bound. This is the first work to rigorously demonstrate performance equivalence between batched and online learning in the nonparametric setting, significantly reducing update frequency while enhancing practical deployability.
This paper studies the generalized linear contextual bandit problem under limited adaptivity, where the number of policy updates is constrained by a budget $M$, aiming to minimize regret. For both fixed- and adaptive-update settings, we propose B-GLinCB—employing batched updates and a novel confidence-region construction—and RS-GLinCB—which incorporates randomization and a tighter, $kappa$-free analysis of the nonlinearity parameter. Our work is the first to eliminate dependence on the nonlinearity constant $kappa$ in generalized linear bandits, achieving a unified $ ilde{O}(sqrt{T})$ regret bound that holds for both stochastic and adversarial context vectors. When $M = Omega(log log T)$, our algorithms attain optimal regret. Notably, RS-GLinCB requires only $ ilde{O}(log^2 T)$ updates, drastically reducing update frequency while preserving theoretical guarantees.
This paper studies dynamic regret minimization in contextual continuum-armed bandits: at each round, an agent selects a decision from a convex action set based on a context, aiming to online optimize a context-dependent objective function. Methodologically, we propose an interior-point algorithm leveraging self-concordant barrier functions, operating under noisy observations. Our contributions are threefold: (i) We establish the first theoretical framework for context-dependent dynamic regret, overcoming the limitations of conventional static regret analysis; (ii) Under Hölder continuity assumptions, we prove that static regret bounds are transferable to the dynamic setting, and our algorithm achieves sublinear dynamic regret; for strongly convex and smooth objectives, it attains the minimax optimal rate up to logarithmic factors; (iii) We rigorously show that sublinear dynamic regret is unattainable when the context–function mapping lacks continuity.
This work addresses the challenge of online recommendation under heterogeneous user preferences, non-stationary context distributions, and the requirement to consistently outperform a baseline policy. The problem is formulated as a linear contextual multi-armed bandit with non-stationary heteroscedastic noise. We propose the first algorithm that simultaneously handles preference heterogeneity, context drift, and baseline constraints by extending the MED strategy to the linear setting, incorporating variance-aware suboptimality gap estimation and a constraint violation control mechanism. Theoretical analysis establishes an instance-dependent regret bound of Õ(κ/Δ̃·d²·log T) and an expected number of constraint violations bounded by Õ(d). Empirical results demonstrate that the proposed method significantly outperforms conservative baselines that ignore either context drift or preference heterogeneity.
This work addresses the problem of achieving optimal decision-making in linear contextual bandits under extremely sparse parameter updates—specifically, only $O(\log\log T)$ times over horizon $T$. The paper proposes two efficient algorithms, BLCE-G and BLCE, both built upon a static scheduling mechanism that is applicable to both small and large action sets and extends naturally to generalized linear models. The key contribution lies in establishing, for the first time, a minimax-optimal regret bound under such infrequent update constraints, up to polylogarithmic factors in $T$. Notably, BLCE eliminates the need for approximate G-optimal design traditionally used in related methods, thereby substantially reducing computational complexity and emerging as the most computationally efficient algorithm to date that attains minimax optimality in this setting.
This work addresses the contextual linear bandit problem where the action set evolves over time according to an exogenous Markov chain, thereby relaxing the conventional i.i.d. context assumption. By constructing a stationary surrogate action set, the non-stationary Markovian contextual problem is reduced to a standard single-context linear bandit. The authors introduce a novel framework that combines delayed updates with a phased learning strategy to control distributional shift, extending the “cheap context” perspective for the first time to Markov-dependent settings. They propose a general reduction framework applicable to uniformly geometrically ergodic chains, enabling online learning of the surrogate mapping even when the transition dynamics are unknown. Under both known and unknown transition distributions, the approach achieves high-probability worst-case regret bounds matching those of the underlying oracle, with only lower-order dependence on the mixing time.
This work addresses the challenge of integrating offline data with online learning in stochastic linear bandits by proposing a novel algorithm that initially leverages offline data to form a prior and subsequently enhances online exploration in an adaptive manner, dynamically balancing the contributions of both sources. Within the structured linear bandit setting, the method is the first to simultaneously outperform purely online and purely offline strategies, achieving a sublinear regret bound with respect to the optimal action. Notably, the regret decreases as the amount of offline data increases. Theoretical analysis establishes a rigorous upper bound on regret, and extensive experiments demonstrate that the proposed approach significantly outperforms existing baselines across various settings.
This work addresses the problem of contextual slate combinatorial online decision-making, where at each round one item is selected from each of $N$ groups and a reward governed by a generalized linear model (GLM) is observed. Under limited adaptivity constraints, the paper proposes two efficient algorithms: a batched variant, B-SlateGLinCB, and a rare-switching variant, RS-SlateGLinCB. Both algorithms rely solely on historical data or infrequent parameter updates, achieving— for the first time in this setting—regret bounds independent of the nonlinearity parameter $\kappa$, while reducing computational complexity to polynomial in $N$, thereby overcoming the exponential blowup of the slate space. Theoretical regret bounds are $O(N d^{3/2} \sqrt{T})$ and $O(N d \sqrt{T})$, respectively. Experiments demonstrate that both methods significantly outperform existing low-adaptivity baselines, closely matching the performance of the fully adaptive Slate-GLM-OFU, and exhibit strong empirical results in large-model context example selection tasks.