design throughput-optimal schedulers

Designs and analyzes scheduling algorithms and control policies that maximize long-run throughput and stabilize queueing systems up to their capacity, producing capacity‑achieving, throughput‑optimal or stability‑optimal schedulers for models that allow preemptive or nonpreemptive service and continuous resource requirements. Builds and integrates mechanisms to enforce rate limits and control operational error rates (for example via admission/rate control or error‑rate feedback) so the scheduler retains throughput and stability guarantees under practical constraints.

designthroughput-optimalschedulers

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.47
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$234K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing scheduling theory struggles to handle multi-resource job scenarios with continuously distributed resource demands, as it relies on the assumption of finitely many job types—a simplification inconsistent with the high heterogeneity observed in real-world workloads. This work proposes the first family of throughput-optimal scheduling policies for continuous multi-resource job models, encompassing both preemptive and non-preemptive variants. The approach employs an adaptive discretization mechanism that dynamically adjusts granularity based on system load and demand distribution. By integrating throughput-optimal control, distribution-aware scheduling, and queueing optimization, the method achieves theoretical optimality while substantially improving computational efficiency. Experiments demonstrate superior performance over state-of-the-art index-based policies under both parametric distributions and real-world Google Borg traces, attaining industry-leading results.

continuous requirement distributionmultiresource-job schedulingqueueing models

Scheduling Servers with Stochastic Bilinear Rewards

Dec 13, 2021
JK
Jung-Hun Kim
🏛️ CREST | ENSAE Paris | London School of Economics

This paper addresses online scheduling in a parallel queue system with multiple job classes and multiple servers, where rewards are unknown, dynamically stochastic, and exhibit a bilinear structure. The objective is to jointly maximize cumulative reward and minimize job holding delay (i.e., holding cost), while ensuring system stability—namely, throughput optimality and bounded queue lengths. We propose the first distributed algorithm integrating three key components: (i) dynamic learning of bilinear bandit rewards, (ii) weighted proportional-fair scheduling, and (iii) marginal-cost correction. Theoretically, the algorithm achieves a sublinear regret bound and guarantees bounded expected queue lengths. Empirically, it significantly outperforms existing baselines in both cumulative reward and average delay across computational service and online platform scenarios.

Balancing reward maximization and fair allocation for system stabilityMinimizing regret while keeping job holding costs boundedScheduling in multi-class parallel-server queues with uncertain rewards

This paper investigates stability and heavy-traffic delay optimality for parallel single-server load balancing systems with heterogeneous service rates, under periodic queue-length observations every $T$ time units; the central dispatcher bases decisions solely on the most recent scaled queue-length ordering and server rates. We propose a general class of scheduling policies that jointly leverage scaled ordering and rate awareness. For the first time, we derive necessary and sufficient conditions for system stability under such policies. Furthermore, we establish sufficient conditions for heavy-traffic delay optimality and prove that, in the heavy-traffic limit, the scaled queue-length vector converges weakly to a deterministic vector multiplied by an exponential random scaling factor. Our analysis integrates stochastic process theory, modeling of periodic information updates, and heavy-traffic scaling limit techniques. This work provides the first rigorous stability criterion and delay optimality guarantee for load balancing in heterogeneous systems operating under limited, periodically updated state information.

Analyzing stability of load balancing with sporadic queue length accessCharacterizing queue length distribution in heavy-traffic asymptotic regimesEstablishing delay optimality conditions for heterogeneous service systems

Improving Multiresource Job Scheduling with Markovian Service Rate Policies

Dec 12, 2024
ZC
Zhongrui Chen
🏛️ University of North Carolina at Chapel Hill | Northwestern University

Scheduling multi-resource jobs—requiring heterogeneous resources such as CPU, memory, and accelerators—in cloud environments poses significant challenges for minimizing average response time under multidimensional resource constraints. Method: This paper introduces the Markovian Service Rate (MSR) policy class, which explicitly models job arrival rates, resource demands, and server capacities. MSR supports diverse system models, including fully preemptive, non-preemptive, and those with setup overheads. Contribution/Results: We prove that MSR policies are throughput-optimal. Moreover, we derive the first tight upper bound on the mean response time—accurate up to an additive constant—thereby establishing a theoretically grounded, practical guideline for scheduler parameter tuning. Empirical evaluation demonstrates substantial improvements in response-time efficiency under realistic multidimensional resource constraints.

Developing simple, analyzable Markovian Service Rate (MSR) policiesEnsuring system stability and tight response time boundsOptimizing multiresource job scheduling to minimize response time

Learning a Discrete Set of Optimal Allocation Rules in a Queueing System with Unknown Service Rate

Feb 04, 2022
SA
Saghar Adler
🏛️ TikTok | University of Iowa | University of Michigan

This paper addresses the optimal admission control problem for an M/M/k/k+N queueing system with unknown service rates, where only arrival epochs and system states (but not service times or departure epochs) are observable, aiming to maximize the long-run average reward. To overcome the exploration-exploitation deadlock inherent in classical certainty-equivalence approaches, we propose a parameterized learning-based self-correcting adaptive control framework that integrates structured policy design, a cautious exploration mechanism, and the certainty-equivalence principle. We establish the first asymptotically optimal learning of extremal heterogeneous optimal policies—namely, “always admit” and “always reject”—and prove that the learned policy converges to the true optimal policy for any underlying service rate. Furthermore, we derive tight finite-time regret upper bounds of both constant and logarithmic order, substantially improving upon existing reinforcement learning methods.

Admission control for M/M/k/k+N queueing systems with unknown service ratesAvoiding suboptimal policies that always block or admit arrivalsMaximizing long-term reward by observing arrivals without service time data

Latest Papers

What's happening recently
View more

This work addresses the challenge of effectively capturing and exploiting short-term scheduling flexibility while guaranteeing long-term worst-case service for multi-task streams. The authors propose a state-based scheduling framework that models worst-case service guarantees as dynamically updatable states and enforces schedulability by constraining state transitions to remain within a well-defined schedulable polytope. For the first time, they fully characterize this schedulable polytope, employ min-plus algebra for an efficient service model representation, and introduce “hyperbolic service” as a novel, dynamically extensible service form. The proposed approach significantly reduces scheduling decision complexity while strictly preserving quality-of-service guarantees, thereby enhancing practical deployability.

min-plus algebraschedulabilityscheduling flexibility

This study addresses the throughput maximization problem for non-preemptive jobs with time windows on both single and multiple machines—a strongly NP-hard scheduling problem. By integrating combinatorial optimization, approximation algorithm design, and pseudo-polynomial time dynamic programming, the authors significantly improve the best-known approximation ratio for the single-machine case from $1.551+\varepsilon$ to $4/3+\varepsilon$, and further refine it to $5/4+\varepsilon$ in pseudo-polynomial time. These results establish the currently best approximation guarantees for this classical scheduling problem and are successfully extended to the multi-machine setting, offering new theoretical insights and algorithmic advances in scheduling under time-window constraints.

Approximation AlgorithmsJob SchedulingNon-Preemptive Scheduling

This study addresses the problem of maximizing throughput in real-time scheduling, focusing on interval scheduling without slack and its generalizations. The authors propose a new model incorporating advance job notifications and establish a constant competitive ratio for proportional weighted throughput under non-preemptive settings. They further demonstrate that, in the absence of bounded processing times, rejection-based preemption cannot guarantee a constant competitive ratio. For instances with at most $k$ distinct processing times, they present a lower bound of $1/(k+1)$ and an algorithm achieving a competitive ratio of $1/(2k)$. These results show that constant competitive ratios are attainable for proportional weights under the advance notification model, yet this property does not extend to general C-/D-benevolent weight functions.

competitive ratiointerval schedulingpreemption

This work addresses the challenge of limited resources in task-specific machine networks that prevent parallel processing of all user jobs. To capture timeliness, the paper introduces "Age of Job" as a novel performance metric and aims to minimize the long-term weighted average job age. By leveraging Lyapunov drift theory, a Max-Weight policy is constructed, and under geometric service times, an optimal Whittle index policy is derived using Whittle index theory. For general service time distributions, a hybrid WIMWF (Whittle Index–Max-Weight Fusion) policy is proposed. Theoretical analysis and simulations demonstrate that WIMWF achieves superior performance under general service time distributions, while the Whittle index policy remains optimal under geometric service times. Moreover, as system scale increases, the NGM policy asymptotically outperforms Max-Weight.

age of jobjob assignmentpreemptive scheduling

Hot Scholars

VN

Vo Nguyen Le Duy

Lecturer at University of Information Technology / Visiting Scientist at RIKEN
Machine LearningData ScienceStatistics
WZ

Wajdi Zaghouani

Associate Professor, Northwestern University
Digital HumanitiesComputational Social SciencesArabic Natural Language ProcessingComputational
NA

Nitin Agarwal

Maulden-Entergy Chair & Donaghey Distinguished Professor, Director, COSMOS Research Center, UALR
Social ComputingData and Web MiningBehavior ModelingInfluence
MM

Mahmoud Meribout

Khalifa University of Science & Technology
Embedded Systems and Instrumentation
TT

Tiejun Tong

Professor of Statistics, Hong Kong Baptist University
StatisticsBiostatisticsMeta-analysisEvidence-based Practice