Score
Designs and analyzes scheduling algorithms and control policies that maximize long-run throughput and stabilize queueing systems up to their capacity, producing capacity‑achieving, throughput‑optimal or stability‑optimal schedulers for models that allow preemptive or nonpreemptive service and continuous resource requirements. Builds and integrates mechanisms to enforce rate limits and control operational error rates (for example via admission/rate control or error‑rate feedback) so the scheduler retains throughput and stability guarantees under practical constraints.
Existing scheduling theory struggles to handle multi-resource job scenarios with continuously distributed resource demands, as it relies on the assumption of finitely many job types—a simplification inconsistent with the high heterogeneity observed in real-world workloads. This work proposes the first family of throughput-optimal scheduling policies for continuous multi-resource job models, encompassing both preemptive and non-preemptive variants. The approach employs an adaptive discretization mechanism that dynamically adjusts granularity based on system load and demand distribution. By integrating throughput-optimal control, distribution-aware scheduling, and queueing optimization, the method achieves theoretical optimality while substantially improving computational efficiency. Experiments demonstrate superior performance over state-of-the-art index-based policies under both parametric distributions and real-world Google Borg traces, attaining industry-leading results.
This paper addresses online scheduling in a parallel queue system with multiple job classes and multiple servers, where rewards are unknown, dynamically stochastic, and exhibit a bilinear structure. The objective is to jointly maximize cumulative reward and minimize job holding delay (i.e., holding cost), while ensuring system stability—namely, throughput optimality and bounded queue lengths. We propose the first distributed algorithm integrating three key components: (i) dynamic learning of bilinear bandit rewards, (ii) weighted proportional-fair scheduling, and (iii) marginal-cost correction. Theoretically, the algorithm achieves a sublinear regret bound and guarantees bounded expected queue lengths. Empirically, it significantly outperforms existing baselines in both cumulative reward and average delay across computational service and online platform scenarios.
This paper investigates stability and heavy-traffic delay optimality for parallel single-server load balancing systems with heterogeneous service rates, under periodic queue-length observations every $T$ time units; the central dispatcher bases decisions solely on the most recent scaled queue-length ordering and server rates. We propose a general class of scheduling policies that jointly leverage scaled ordering and rate awareness. For the first time, we derive necessary and sufficient conditions for system stability under such policies. Furthermore, we establish sufficient conditions for heavy-traffic delay optimality and prove that, in the heavy-traffic limit, the scaled queue-length vector converges weakly to a deterministic vector multiplied by an exponential random scaling factor. Our analysis integrates stochastic process theory, modeling of periodic information updates, and heavy-traffic scaling limit techniques. This work provides the first rigorous stability criterion and delay optimality guarantee for load balancing in heterogeneous systems operating under limited, periodically updated state information.
Scheduling multi-resource jobs—requiring heterogeneous resources such as CPU, memory, and accelerators—in cloud environments poses significant challenges for minimizing average response time under multidimensional resource constraints. Method: This paper introduces the Markovian Service Rate (MSR) policy class, which explicitly models job arrival rates, resource demands, and server capacities. MSR supports diverse system models, including fully preemptive, non-preemptive, and those with setup overheads. Contribution/Results: We prove that MSR policies are throughput-optimal. Moreover, we derive the first tight upper bound on the mean response time—accurate up to an additive constant—thereby establishing a theoretically grounded, practical guideline for scheduler parameter tuning. Empirical evaluation demonstrates substantial improvements in response-time efficiency under realistic multidimensional resource constraints.
This paper addresses the optimal admission control problem for an M/M/k/k+N queueing system with unknown service rates, where only arrival epochs and system states (but not service times or departure epochs) are observable, aiming to maximize the long-run average reward. To overcome the exploration-exploitation deadlock inherent in classical certainty-equivalence approaches, we propose a parameterized learning-based self-correcting adaptive control framework that integrates structured policy design, a cautious exploration mechanism, and the certainty-equivalence principle. We establish the first asymptotically optimal learning of extremal heterogeneous optimal policies—namely, “always admit” and “always reject”—and prove that the learned policy converges to the true optimal policy for any underlying service rate. Furthermore, we derive tight finite-time regret upper bounds of both constant and logarithmic order, substantially improving upon existing reinforcement learning methods.
This work addresses the challenge of effectively capturing and exploiting short-term scheduling flexibility while guaranteeing long-term worst-case service for multi-task streams. The authors propose a state-based scheduling framework that models worst-case service guarantees as dynamically updatable states and enforces schedulability by constraining state transitions to remain within a well-defined schedulable polytope. For the first time, they fully characterize this schedulable polytope, employ min-plus algebra for an efficient service model representation, and introduce “hyperbolic service” as a novel, dynamically extensible service form. The proposed approach significantly reduces scheduling decision complexity while strictly preserving quality-of-service guarantees, thereby enhancing practical deployability.
This study addresses the throughput maximization problem for non-preemptive jobs with time windows on both single and multiple machines—a strongly NP-hard scheduling problem. By integrating combinatorial optimization, approximation algorithm design, and pseudo-polynomial time dynamic programming, the authors significantly improve the best-known approximation ratio for the single-machine case from $1.551+\varepsilon$ to $4/3+\varepsilon$, and further refine it to $5/4+\varepsilon$ in pseudo-polynomial time. These results establish the currently best approximation guarantees for this classical scheduling problem and are successfully extended to the multi-machine setting, offering new theoretical insights and algorithmic advances in scheduling under time-window constraints.
This study addresses the problem of maximizing throughput in real-time scheduling, focusing on interval scheduling without slack and its generalizations. The authors propose a new model incorporating advance job notifications and establish a constant competitive ratio for proportional weighted throughput under non-preemptive settings. They further demonstrate that, in the absence of bounded processing times, rejection-based preemption cannot guarantee a constant competitive ratio. For instances with at most $k$ distinct processing times, they present a lower bound of $1/(k+1)$ and an algorithm achieving a competitive ratio of $1/(2k)$. These results show that constant competitive ratios are attainable for proportional weights under the advance notification model, yet this property does not extend to general C-/D-benevolent weight functions.
This work addresses the challenge of limited resources in task-specific machine networks that prevent parallel processing of all user jobs. To capture timeliness, the paper introduces "Age of Job" as a novel performance metric and aims to minimize the long-term weighted average job age. By leveraging Lyapunov drift theory, a Max-Weight policy is constructed, and under geometric service times, an optimal Whittle index policy is derived using Whittle index theory. For general service time distributions, a hybrid WIMWF (Whittle Index–Max-Weight Fusion) policy is proposed. Theoretical analysis and simulations demonstrate that WIMWF achieves superior performance under general service time distributions, while the Whittle index policy remains optimal under geometric service times. Moreover, as system scale increases, the NGM policy asymptotically outperforms Max-Weight.