contextual bandits

Design, build, and evaluate online decision-making algorithms and policies that select actions given observed context to maximize cumulative reward, covering multi-armed, contextual, linear, and slate bandit formulations and their stochastic and delayed-feedback variants. Work includes constructing batched, rarely-switching, and limited-adaptivity procedures, formulating experimental-design-based exploration, implementing efficient per-round policy computation, and proving sublinear regret and other performance guarantees under streaming or delayed feedback.

contextualbandits

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.69
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$209K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of high-probability optimal regret bounds for policy optimization methods in stochastic contextual multi-armed bandits (CMAB). To bridge the gap between theoretical guarantees and practical applicability, the paper proposes an efficient algorithm that integrates policy optimization with general offline function approximation. The method establishes, for the first time, a high-probability optimal regret bound of $\widetilde{O}(\sqrt{K|\mathcal{A}| \log|\mathcal{F}|})$ for policy optimization in CMAB, where $K$ denotes the number of contexts, $|\mathcal{A}|$ the number of actions, and $|\mathcal{F}|$ the complexity of the policy class. Both theoretical analysis and empirical experiments corroborate the algorithm’s effectiveness and optimality, thereby providing a solid foundation for deploying policy optimization in real-world contextual bandit settings.

contextual banditsfunction approximationpolicy optimization

This paper addresses the problem of safely and efficiently improving online learning performance for stochastic multi-armed bandits (MAB) and combinatorial MAB under offline–online distribution shift. To tackle distribution mismatch between offline and online data, we propose MIN-UCB—a novel adaptive algorithm that automatically discards low-quality offline data when the shift magnitude is unknown, ensuring safety, and actively leverages bounded-shift offline information to substantially reduce regret when the shift is bounded. We establish, for the first time, a theoretical lower bound on regret for MAB with biased offline data. We prove that MIN-UCB achieves tight instance-dependent and instance-independent regret bounds—strictly outperforming classical UCB. Both theoretical analysis and numerical experiments demonstrate its robustness and superiority across diverse shift regimes.

Addressing distribution mismatch between offline data and online rewards.Developing adaptive policies that selectively use informative offline data.Leveraging offline data to enhance online learning in bandit problems.

Batched Nonparametric Contextual Bandits

Feb 27, 2024
RJ
Rong Jiang
🏛️ University of Chicago

This paper studies the nonparametric contextual bandit problem under batch constraints, where the reward function is an unknown smooth function of covariates and the policy is updated only upon completion of each batch. To address the exploration–exploitation trade-off inherent in batched learning, we propose a dynamic binning mechanism: bin widths adaptively scale with batch sizes, integrated with nonparametric regression and minimax analysis to achieve efficient estimation within the batched learning framework. We theoretically establish that only a constant number of policy updates suffice to attain the optimal online regret bound—up to logarithmic factors—and provide a matching lower bound. This is the first work to rigorously demonstrate performance equivalence between batched and online learning in the nonparametric setting, significantly reducing update frequency while enhancing practical deployability.

Dynamic covariate space splitting for optimal regretEstablish minimax regret lower bound and optimal algorithmStudy nonparametric contextual bandits with batch constraints

Generalized Linear Bandits with Limited Adaptivity

Apr 10, 2024
AS
Ayush Sawarni
🏛️ Microsoft Research | Indian Institute of Science

This paper studies the generalized linear contextual bandit problem under limited adaptivity, where the number of policy updates is constrained by a budget $M$, aiming to minimize regret. For both fixed- and adaptive-update settings, we propose B-GLinCB—employing batched updates and a novel confidence-region construction—and RS-GLinCB—which incorporates randomization and a tighter, $kappa$-free analysis of the nonlinearity parameter. Our work is the first to eliminate dependence on the nonlinearity constant $kappa$ in generalized linear bandits, achieving a unified $ ilde{O}(sqrt{T})$ regret bound that holds for both stochastic and adversarial context vectors. When $M = Omega(log log T)$, our algorithms attain optimal regret. Notably, RS-GLinCB requires only $ ilde{O}(log^2 T)$ updates, drastically reducing update frequency while preserving theoretical guarantees.

Algorithms for fixed and adaptive policy update settings.Eliminating dependence on key non-linearity parameter.Generalized linear contextual bandits with limited adaptivity.

Contextual Continuum Bandits: Static Versus Dynamic Regret

Jun 09, 2024
AA
A. Akhavan
🏛️ Istituto Italiano di Tecnologia | Ecole Polytechnique | University College London | ENS AE | IP Paris

This paper studies dynamic regret minimization in contextual continuum-armed bandits: at each round, an agent selects a decision from a convex action set based on a context, aiming to online optimize a context-dependent objective function. Methodologically, we propose an interior-point algorithm leveraging self-concordant barrier functions, operating under noisy observations. Our contributions are threefold: (i) We establish the first theoretical framework for context-dependent dynamic regret, overcoming the limitations of conventional static regret analysis; (ii) Under Hölder continuity assumptions, we prove that static regret bounds are transferable to the dynamic setting, and our algorithm achieves sublinear dynamic regret; for strongly convex and smooth objectives, it attains the minimax optimal rate up to logarithmic factors; (iii) We rigorously show that sublinear dynamic regret is unattainable when the context–function mapping lacks continuity.

Extending static regret algorithms to achieve sub-linear dynamic regret.Minimizing dynamic regret in contextual continuum bandits.Proposing an algorithm for sub-linear dynamic regret with noisy observations.

Latest Papers

What's happening recently
View more

This work addresses the challenge of online recommendation under heterogeneous user preferences, non-stationary context distributions, and the requirement to consistently outperform a baseline policy. The problem is formulated as a linear contextual multi-armed bandit with non-stationary heteroscedastic noise. We propose the first algorithm that simultaneously handles preference heterogeneity, context drift, and baseline constraints by extending the MED strategy to the linear setting, incorporating variance-aware suboptimality gap estimation and a constraint violation control mechanism. Theoretical analysis establishes an instance-dependent regret bound of Õ(κ/Δ̃·d²·log T) and an expected number of constraint violations bounded by Õ(d). Empirical results demonstrate that the proposed method significantly outperforms conservative baselines that ignore either context drift or preference heterogeneity.

conservative constraintcontext driftcontextual bandits

This work addresses the problem of achieving optimal decision-making in linear contextual bandits under extremely sparse parameter updates—specifically, only $O(\log\log T)$ times over horizon $T$. The paper proposes two efficient algorithms, BLCE-G and BLCE, both built upon a static scheduling mechanism that is applicable to both small and large action sets and extends naturally to generalized linear models. The key contribution lies in establishing, for the first time, a minimax-optimal regret bound under such infrequent update constraints, up to polylogarithmic factors in $T$. Notably, BLCE eliminates the need for approximate G-optimal design traditionally used in related methods, thereby substantially reducing computational complexity and emerging as the most computationally efficient algorithm to date that attains minimax optimality in this setting.

batched learningcomputational efficiencylinear contextual bandits

This work addresses the contextual linear bandit problem where the action set evolves over time according to an exogenous Markov chain, thereby relaxing the conventional i.i.d. context assumption. By constructing a stationary surrogate action set, the non-stationary Markovian contextual problem is reduced to a standard single-context linear bandit. The authors introduce a novel framework that combines delayed updates with a phased learning strategy to control distributional shift, extending the “cheap context” perspective for the first time to Markov-dependent settings. They propose a general reduction framework applicable to uniformly geometrically ergodic chains, enabling online learning of the surrogate mapping even when the transition dynamics are unknown. Under both known and unknown transition distributions, the approach achieves high-probability worst-case regret bounds matching those of the underlying oracle, with only lower-order dependence on the mixing time.

contextual banditslinear banditsMarkovian contextual bandits

This work addresses the challenge of integrating offline data with online learning in stochastic linear bandits by proposing a novel algorithm that initially leverages offline data to form a prior and subsequently enhances online exploration in an adaptive manner, dynamically balancing the contributions of both sources. Within the structured linear bandit setting, the method is the first to simultaneously outperform purely online and purely offline strategies, achieving a sublinear regret bound with respect to the optimal action. Notably, the regret decreases as the amount of offline data increases. Theoretical analysis establishes a rigorous upper bound on regret, and extensive experiments demonstrate that the proposed approach significantly outperforms existing baselines across various settings.

linear banditsoffline datasetoffline-to-online learning

This work addresses the problem of contextual slate combinatorial online decision-making, where at each round one item is selected from each of $N$ groups and a reward governed by a generalized linear model (GLM) is observed. Under limited adaptivity constraints, the paper proposes two efficient algorithms: a batched variant, B-SlateGLinCB, and a rare-switching variant, RS-SlateGLinCB. Both algorithms rely solely on historical data or infrequent parameter updates, achieving— for the first time in this setting—regret bounds independent of the nonlinearity parameter $\kappa$, while reducing computational complexity to polynomial in $N$, thereby overcoming the exponential blowup of the slate space. Theoretical regret bounds are $O(N d^{3/2} \sqrt{T})$ and $O(N d \sqrt{T})$, respectively. Experiments demonstrate that both methods significantly outperform existing low-adaptivity baselines, closely matching the performance of the fully adaptive Slate-GLM-OFU, and exhibit strong empirical results in large-model context example selection tasks.

combinatorial selectioncontextual slate banditsgeneralized linear models

Hot Scholars

MV

Michal Valko

Chief Models Officer @ Stealth Startup, Inria & MVA - Ex: Llama at Meta; Gemini and BYOL @ Deepmind
large language modelsreasoningfine-tuningtest-time computation
ZD

Zhongxiang Dai

Assistant Professor, The Chinese University of Hong Kong, Shenzhen
Machine LearningData-Centric AILarge Language ModelsMulti-Armed Bandits
VP

Vianney Perchet

Crest, ENSAE & Criteo AI Lab
Game TheoryMulti-armed BanditMachine Learning
MH

Min-hwan Oh

Seoul National University
Reinforcement LearningBandit AlgorithmsMachine Learning