contextual bandit scheduling

Designs and implements online decision systems that schedule and select actions or behaviors using contextual multi-armed bandit models, including specifying context representations, reward signals, and exploration–exploitation policies. Builds behavior-bandit schedulers that adaptively choose among candidate policies or behaviors and analyzes their empirical performance (e.g., regret or cumulative reward) under changing contexts.

contextualbanditscheduling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Contextual Bandits with Arm Request Costs and Delays

Oct 17, 2024
LW
Lai Wei
🏛️ University of Michigan

This paper studies the delay-aware contextual bandit problem: a learner dynamically selects a subset of arms under stochastic contexts, where each selection incurs a random delay and a context-dependent arm request cost; the total time horizon is determined by the cumulative delays of selected subsets. The problem is formulated as a semi-Markov decision process (SMDP), the first to jointly model context dependence, arm-specific request costs, and stochastic delays. We propose an online algorithm derived from the Bellman optimality equation and establish a tight $O(sqrt{T})$ regret bound under a realizability assumption—matching the optimal rate of classical contextual bandits. The theoretical analysis is rigorous, and extensive experiments on synthetic benchmarks and a movie recommendation dataset validate both the algorithm’s empirical effectiveness and its theoretical guarantees.

Balances exploration-exploitation trade-off while minimizing cumulative regretExtends contextual bandits to handle action delays and set switchingOptimizes arm selection under latency constraints from unknown distributions

Neural Exploitation and Exploration of Contextual Bandits

May 05, 2023
YB
Yikun Ban
🏛️ University of Illinois at Urbana-Champaign

This paper addresses the exploration-exploitation trade-off in contextual multi-armed bandits. We propose EE-Net, a dual-neural-network architecture: one network models the reward function for efficient exploitation, while the other directly learns instance-dependent exploration gains—bypassing conventional statistical confidence bounds—to enable adaptive exploration. To our knowledge, this is the first work to explicitly model exploration gains using neural networks. Theoretical analysis establishes an instance-dependent regret upper bound of $ ilde{O}(sqrt{T})$. Empirical evaluation on multiple real-world datasets demonstrates that EE-Net significantly outperforms both linear and state-of-the-art neural contextual bandit baselines, validating its modeling flexibility and generalization capability.

Achieves sublinear regret and outperforms existing baselines on real datasetsProposes EE-Net for neural contextual bandit exploitation and explorationUses separate neural networks to adaptively learn reward and exploration gains

Contextual Bandits with Stage-wise Constraints

Jan 15, 2024
AP
Aldo Pacchiano
🏛️ Boston University | Broad Institute of MIT and Harvard | Amazon | University of California Berkeley

This paper studies contextual bandits with stage-wise constraints, requiring each decision to satisfy constraints simultaneously under both high-probability and expectation-based feasibility criteria, while maximizing cumulative reward and ensuring real-time constraint satisfaction. We propose, for the first time, a differentiated scaling mechanism that models confidence set radii separately for rewards and costs. A unified framework is developed to handle single or multiple constraints—whether linear or nonlinear in structure. Leveraging UCB-style exploration, eluder dimension analysis, and joint dual-constraint modeling, we establish an optimal $ ilde{O}(sqrt{T})$ regret bound and provide matching upper and lower bounds. The theoretical analysis is rigorous, and empirical simulations consistently validate the theoretical guarantees. The algorithm is scalable and applicable to complex, nonlinear constraint settings.

Addressing contextual bandits with stage-wise constraintsEnsuring constraints are satisfied both probabilistically and expectedlyExtending solutions from linear to non-linear reward-cost functions

Contextual Combinatorial Bandits with Changing Action Sets via Gaussian Processes

Oct 05, 2021
AN
Andi Nika
🏛️ Max Planck Institute for Software Systems | EPFL | Bilkent University

This paper studies the contextual combinatorial bandit problem with a dynamically evolving base arm set over time, aiming to maximize cumulative reward. To address the dual challenges of time-varying feasible action sets and context-dependent rewards, we introduce Gaussian process (GP) modeling into this framework for the first time, proposing the O’CLOK-UCB algorithm and its sparse GP-accelerated variant. Our method integrates kernelized UCB, combinatorial feasibility constraints, and Lipschitz continuity analysis. We establish a sublinear regret bound of Õ(√(λ∗(K)KTγ_T)), where λ∗(K) is the largest eigenvalue of the action covariance matrix and γ_T is the maximum information gain—revealing their coupled impact on regret. Empirical evaluation on real-world datasets demonstrates significant improvements over existing UCB-based approaches, confirming both theoretical rigor and practical efficacy.

Modeling base arm outcomes via Gaussian Process contextual dependenciesOptimizing cumulative reward under changing action set constraintsSolving combinatorial bandits with time-varying available actions

Batched Nonparametric Contextual Bandits

Feb 27, 2024
RJ
Rong Jiang
🏛️ University of Chicago

This paper studies the nonparametric contextual bandit problem under batch constraints, where the reward function is an unknown smooth function of covariates and the policy is updated only upon completion of each batch. To address the exploration–exploitation trade-off inherent in batched learning, we propose a dynamic binning mechanism: bin widths adaptively scale with batch sizes, integrated with nonparametric regression and minimax analysis to achieve efficient estimation within the batched learning framework. We theoretically establish that only a constant number of policy updates suffice to attain the optimal online regret bound—up to logarithmic factors—and provide a matching lower bound. This is the first work to rigorously demonstrate performance equivalence between batched and online learning in the nonparametric setting, significantly reducing update frequency while enhancing practical deployability.

Dynamic covariate space splitting for optimal regretEstablish minimax regret lower bound and optimal algorithmStudy nonparametric contextual bandits with batch constraints

Latest Papers

What's happening recently
View more

This work proposes a generalized multi-armed bandit and stopping problem framework grounded in behavioral preferences, departing from conventional modeling paradigms that rely on predefined states, rewards, and transition dynamics. Starting from the decision maker’s preferences over local temporal plans, the authors develop a generalized stopping representation through behavioral axioms and introduce a calendar-time cross-plan pricing mechanism. Under compact time constraints, this approach yields a rested-bandit model exhibiting index optimality. The key innovation lies in interpreting the index as the shadow price of advancing a local clock, thereby unifying—within a single preference-based framework—a diverse array of decision models, including expected utility, learning, robust, rank-dependent, Choquet, and Pandora’s box formulations. This constitutes the first theoretical foundation for index policies rooted entirely in preference theory.

Behavioral PreferencesDecision TheoryMulti-Armed Bandits

This work addresses the challenge of online recommendation under heterogeneous user preferences, non-stationary context distributions, and the requirement to consistently outperform a baseline policy. The problem is formulated as a linear contextual multi-armed bandit with non-stationary heteroscedastic noise. We propose the first algorithm that simultaneously handles preference heterogeneity, context drift, and baseline constraints by extending the MED strategy to the linear setting, incorporating variance-aware suboptimality gap estimation and a constraint violation control mechanism. Theoretical analysis establishes an instance-dependent regret bound of Õ(κ/Δ̃·d²·log T) and an expected number of constraint violations bounded by Õ(d). Empirical results demonstrate that the proposed method significantly outperforms conservative baselines that ignore either context drift or preference heterogeneity.

conservative constraintcontext driftcontextual bandits

This work addresses the problem of achieving optimal decision-making in linear contextual bandits under extremely sparse parameter updates—specifically, only $O(\log\log T)$ times over horizon $T$. The paper proposes two efficient algorithms, BLCE-G and BLCE, both built upon a static scheduling mechanism that is applicable to both small and large action sets and extends naturally to generalized linear models. The key contribution lies in establishing, for the first time, a minimax-optimal regret bound under such infrequent update constraints, up to polylogarithmic factors in $T$. Notably, BLCE eliminates the need for approximate G-optimal design traditionally used in related methods, thereby substantially reducing computational complexity and emerging as the most computationally efficient algorithm to date that attains minimax optimality in this setting.

batched learningcomputational efficiencylinear contextual bandits

This work addresses the problem of identifying an $\varepsilon$-optimal policy in stochastic contextual bandits with $s$-sparse rewards. The authors propose an algorithm whose sample complexity depends only on the sparsity level $s$, rather than high-degree polynomials of the action space size $|A|$. By integrating information-theoretic analysis based on the Decision-Estimation Coefficient (DEC) with low-variance exploration techniques, the method applies to policy classes with bounded Natarajan dimension and extends to combinatorial semi-bandit settings. The resulting sample complexity is $\widetilde{O}\big((s/\varepsilon^2 + |A|/\varepsilon) \log(|\Pi|/\delta)\big)$, which is near-optimal up to logarithmic factors. This significantly improves upon prior results that scaled with $|A|^9$ and provides the first tight upper bound featuring no higher-order dependence on $|A|$.

contextual banditsmulticlass classificationsample complexity

Hot Scholars

HY

Haizhao Yang

Department of Mathematics, Department of Computer Science, University of Maryland College Park
Data sciencemachine learninghigh-performance computingnumerical linear algebra
DZ

Debin Zhao

Dept. of Computer Science,Harbin Institute of Technology
Video codingImage and Video ProcessingData Compression
XZ

Xiantao Zhang

Beihang University
Large Language ModelsNatural Language ProcessingArtificial IntelligenceData Curation