Contextual Bandits with Arm Request Costs and Delays

📅 2024-10-17
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper studies the delay-aware contextual bandit problem: a learner dynamically selects a subset of arms under stochastic contexts, where each selection incurs a random delay and a context-dependent arm request cost; the total time horizon is determined by the cumulative delays of selected subsets. The problem is formulated as a semi-Markov decision process (SMDP), the first to jointly model context dependence, arm-specific request costs, and stochastic delays. We propose an online algorithm derived from the Bellman optimality equation and establish a tight $O(sqrt{T})$ regret bound under a realizability assumption—matching the optimal rate of classical contextual bandits. The theoretical analysis is rigorous, and extensive experiments on synthetic benchmarks and a movie recommendation dataset validate both the algorithm’s empirical effectiveness and its theoretical guarantees.

Technology Category

Machine Learning: Online Learning & BanditsSearch and Optimization: Learning to SearchReasoning under Uncertainty: Stochastic Optimization

Application Category

Responsible Web: Human-perceived consequences of algorithmic deployment on the webSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphs
📝 Abstract
We introduce a novel extension of the contextual bandit problem, where new sets of arms can be requested with stochastic time delays and associated costs. In this setting, the learner can select multiple arms from a decision set, with each selection taking one unit of time. The problem is framed as a special case of semi-Markov decision processes (SMDPs). The arm contexts, request times, and costs are assumed to follow an unknown distribution. We consider the regret of an online learning algorithm with respect to the optimal policy that achieves the maximum average reward. By leveraging the Bellman optimality equation, we design algorithms that can effectively select arms and determine the appropriate time to request new arms, thereby minimizing their regret. Under the realizability assumption, we analyze the proposed algorithms and demonstrate that their regret upper bounds align with established results in the contextual bandit literature. We validate the algorithms through experiments on simulated data and a movie recommendation dataset, showing that their performance is consistent with theoretical analyses.
Problem

Research questions and friction points this paper is trying to address.

Extends contextual bandits to handle action delays and set switching
Optimizes arm selection under latency constraints from unknown distributions
Balances exploration-exploitation trade-off while minimizing cumulative regret
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latency-aware contextual bandit framework with action delays
Contextual online arm filtering algorithm balancing exploration
Application to cryo-EM data collection maximizing cumulative reward
🔎 Similar Papers
2024-07-24arXiv.orgCitations: 4
💼 Related Jobs
No related jobs found.
University of Michigan
L
Lai Wei
Life Sciences Institute, University of Michigan
A
Ambuj Tewari
Department of Statistics, University of Michigan
M
M. Cianfrocco
Life Sciences Institute, University of Michigan