Learning While Scheduling Jobs under Context-Dependent Service Rates: An Anytime Rate-Optimal Algorithm

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of suboptimal decay and reliance on a known time horizon in scheduling with unknown service rates within contextual queueing bandits. To overcome these challenges, this work proposes the WISE algorithm, which leverages a capacity slackness assumption and integrates confidence interval elimination, workload potential function analysis, and elliptical potential counting theory to achieve optimal scheduling without prior knowledge of the time horizon. The primary contribution is the first derivation of a contextual queueing bandit (CQB) lower bound to quantify the impact of capacity slackness on regret, alongside a relaxation of the covariance lower bound assumption. Theoretically, the proposed method attains a rate-optimal queue-length regret of $\widetilde{O}(t^{-1/2})$ at any time $t$, while simulations demonstrate its ability to maintain low regret even when capacity slackness fails.
📝 Abstract
We study contextual queueing bandits, where a learner schedules jobs while learning unknown service rates modeled by logistic functions of job-server features. Performance is measured by queue length regret, the expected excess queue length at round $t$ relative to an oracle that knows the service rates. Existing decaying-regret guarantees either have a suboptimal decay rate or require a known fixed horizon. They also assume context-wise slack and a strictly positive minimum eigenvalue of the feature covariance. In this paper, we propose WISE (Widest Interval Selection with Elimination), achieving rate-optimal $\widetilde{\mathcal O}(t^{-1/2})$ queue length regret at every sufficiently large time without knowing the horizon. We assume capacity slack, meaning that expected incoming workload under best-server service is below service capacity, and impose no covariance lower bound. Our analysis uses a workload potential measuring the expected service attempts needed by waiting jobs on their best servers. Its drift on nonempty rounds combines a negative term ensured by capacity slack with errors from suboptimal service choices. Then an elliptical potential count bounds how often WISE selects wide confidence intervals, thereby limiting the number of rounds with large service errors. We also sharpen the arrival-rate dependence of an existing lower bound and make its dependence on feature dimension and server count explicit. We prove another lower bound that quantifies the increase in regret as the normalized capacity slack decreases; to our knowledge, this is the first such lower bound for CQB. Simulations show small regret even when context-wise slack fails.
Problem

Research questions and friction points this paper is trying to address.

contextual queueing bandits
queue length regret
job scheduling
service rate learning
🔎 Similar Papers
No similar papers found.