🤖 AI Summary
This study addresses the problem of learning service rates to minimize queue-length regret in single-server non-preemptive queuing systems with unknown context distributions. It proposes the Learn-Clear-Plan (LCP) algorithmic framework, which integrates logistic regression for parameter estimation with finite-horizon Bellman recursion for online decision-making, and establishes the key property that the optimal action under a known model depends on the remaining horizon. The primary contribution lies in deriving theoretical bounds for non-preemptive scheduling, achieving an $\tilde{O}(\sqrt{d/T})$ queue-length regret that matches the $\Omega(\min\{1/\sqrt{d}, \sqrt{d/T}\})$ lower bound. Furthermore, this work provides Shortest Expected Processing Time (SEPT) tracking error guarantees in settings where the time horizon is unknown.
📝 Abstract
We study nonpreemptive contextual queueing bandits in a single-server system. Each job is represented by a $d$-dimensional context vector; in each round, a job may arrive with its context drawn from an unknown distribution $\mathcal{D}$, and its departure probability is determined by a logistic model of that context vector with an unknown parameter $θ^*$. The server learns from service outcomes while deciding which waiting job to serve and whether to idle, aiming to minimize queue-length regret, the gap between its expected terminal queue length and the minimum achievable by an admissible policy. Once selected, a job must be served until completion, and we refer to this as the nonpreemptive setting. A central challenge is that, even with full model knowledge, the optimal policy cannot in general be characterized by a simple myopic rule, since the optimal action can change with the remaining horizon at the same queue state. Nevertheless, when the model and horizon are known, the optimal action can be obtained through a finite-horizon Bellman recursion. Motivated by this, we propose Learn--Clear--Plan (LCP), which estimates the system and uses the resulting Bellman recursion to make horizon-dependent decisions. LCP achieves $\widetilde{O}(\sqrt{d/T})$ queue-length regret, while a lower-bound construction gives $Ω(\min\{1/\sqrt{d},\sqrt{d/T}\})$ regret for every learning policy on some instance, establishing optimality up to polylogarithmic factors when $T\ge d^2$. When the horizon is unknown, no horizon-independent policy achieves vanishing regret against the finite-horizon optimum. We therefore use SEPT, the policy that serves a waiting job with the highest probability of departure, as a fixed reference, and suggest an estimated-SEPT algorithm that achieves a tracking error of $\widetilde{O}(\sqrt{d/t})$ without knowing the model.