build job schedulers

Designs, implements, and analyzes schedulers and scheduling systems that assign jobs or tasks to execution resources over time, including runtime/request scheduling, pipeline coordination, priority and fairness handling, multi-region and elastic scaling, and integrated workload placement. Develops and evaluates scheduling algorithms and optimizations (priority scheduling, scheduling algorithm design, throughput/latency objectives), and performs scheduler tuning, performance optimization, and implementation of scheduling infrastructure and controls.

buildjobschedulers

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Workload Schedulers -- Genesis, Algorithms and Differences

Nov 13, 2025
LS
L. Sliwko
🏛️ University of Westminster

This paper addresses the lack of clarity regarding the diversity and evolutionary trajectories of modern workload schedulers. We propose a cross-layer taxonomy comprising three categories: OS process scheduling, cluster job scheduling, and big-data scheduling. Through algorithmic feature analysis and historical comparative study, we systematically characterize the design rationales, optimization objectives, and technological evolution of these schedulers, uncovering shared design patterns across local and distributed environments. Our key contribution is the first unified classification framework, which identifies three fundamental differentiating dimensions: resource abstraction granularity, scheduling timing, and feedback mechanism. Based on this analysis, we distill general-purpose scheduling design principles targeting heterogeneity, scalability, and QoS guarantees. The study provides both theoretical foundations and practical guidance for scheduler selection, cross-layer coordination optimization, and next-generation scheduler architecture design.

Analyzing scheduler evolution from early adoptions to modern implementationsCategorizing modern workload schedulers into three distinct classesComparing scheduling strategies across local and distributed systems

Is the GPU Half-Empty or Half-Full? Practical Scheduling Techniques for LLMs

Oct 23, 2024
FK
Ferdinand Kossmann
🏛️ MIT | Databricks

This work addresses inefficient GPU resource scheduling in large language model (LLM) inference serving. We propose a two-tier cooperative scheduling framework: server-level scheduling for load balancing and service-level scheduling optimized for request latency sensitivity. Our approach introduces a lightweight, deployable dynamic priority queue and a preemptive batching mechanism—requiring no modifications to models, hardware, or underlying inference frameworks—and maintains full compatibility with mainstream LLM serving systems. Evaluated under real production workloads, it reduces average tail latency by 22%, improves GPU utilization by 18%, and increases throughput by 15% over state-of-the-art production-grade scheduling policies. The core contribution is a practical, high-performance scheduling paradigm that achieves significant efficiency gains with minimal implementation overhead, delivering a production-ready resource optimization solution for LLM inference serving.

GPU OptimizationLarge Language ModelsResource Allocation

This study investigates the scalability and performance of process and thread schedulers under memory-intensive workloads in multi-core shared-memory systems, focusing on a 3D tensor row-sorting task. The authors design and evaluate several scheduling strategies: on the thread side, an AIMD-based adaptive chunking mechanism inspired by TCP congestion control is introduced, coupled with exponential weighted moving average to dynamically adjust concurrency; on the process side, a bounded prolific/collective model is employed alongside one-to-one, one-to-many, and many-to-many pipelined communication patterns to enable flexible task distribution. Experimental results on a 24-core x86-64 platform demonstrate that thread-level scheduling consistently outperforms process-level scheduling, with dynamic and guided strategies achieving the best performance, while the many-to-many pipeline exhibits superior scalability for large-scale tasks.

many-core systemsprocess-based schedulingscalability

A Reinforcement Learning-Driven Task Scheduling Algorithm for Multi-Tenant Distributed Systems

Aug 11, 2025
XZ
Xiaopei Zhang
🏛️ University of California, Los Angeles | Institute of Automation, Chinese Academy of Sciences | University of the Chinese Academy of Sciences

Addressing the challenge of jointly managing dynamic resource fluctuations, heterogeneous tenant requirements, and fairness guarantees in multi-tenant distributed systems, this paper proposes a reinforcement learning–based adaptive task scheduling framework. We formulate scheduling as a Markov decision process and design a multi-objective reward function integrating task latency, resource utilization, and tenant fairness. Using the Proximal Policy Optimization (PPO) algorithm, we jointly train policy and value networks to enhance training stability and cross-scenario generalization. Evaluated on real-data-driven multi-tenant workloads, our approach significantly outperforms state-of-the-art schedulers: it reduces average task latency by 23.6%, improves resource utilization by 18.4%, and increases the Jain’s fairness index by 0.31 across tenants. The framework balances theoretical rigor with practical deployability, offering a principled yet scalable solution for fair and efficient resource orchestration in dynamic multi-tenant environments.

Balancing latency, resource use, and tenant fairnessDynamic task scheduling in multi-tenant distributed systemsReinforcement learning for adaptive scheduling decisions

Timing Analysis and Priority-driven Enhancements of ROS 2 Multi-threaded Executors

May 01, 2023
HS
Hoora Sobhani
🏛️ University of California, Riverside | San Diego State University

To address the lack of systematic response-time analysis and real-time scheduling guarantees in ROS 2’s multi-threaded executor, this paper introduces the first response-time analysis framework tailored to its kernel-level execution semantics. The framework supports modeling of arbitrary- and constrained-deadline task chains and precisely captures mutual exclusion among callback groups. We further propose a priority-driven scheduling enhancement mechanism that optimizes critical-path response times while preserving schedulability. Experimental evaluation on the Jetson AGX Xavier platform demonstrates that our framework yields tighter safe upper bounds on response time, reduces average response time of critical chains by a significant margin, and improves overall system schedulability by 23.6%.

Multi-threaded ExecutorROS 2Task Scheduling

Latest Papers

What's happening recently
View more

Existing scheduling theory struggles to handle multi-resource job scenarios with continuously distributed resource demands, as it relies on the assumption of finitely many job types—a simplification inconsistent with the high heterogeneity observed in real-world workloads. This work proposes the first family of throughput-optimal scheduling policies for continuous multi-resource job models, encompassing both preemptive and non-preemptive variants. The approach employs an adaptive discretization mechanism that dynamically adjusts granularity based on system load and demand distribution. By integrating throughput-optimal control, distribution-aware scheduling, and queueing optimization, the method achieves theoretical optimality while substantially improving computational efficiency. Experiments demonstrate superior performance over state-of-the-art index-based policies under both parametric distributions and real-world Google Borg traces, attaining industry-leading results.

continuous requirement distributionmultiresource-job schedulingqueueing models

A Real-Time Digital Twin for Adaptive Scheduling

Dec 21, 2025
YZ
Yihe Zhang
🏛️ University of Illinois Chicago | Argonne National Laboratory

HPC workloads are becoming increasingly heterogeneous, rendering traditional static heuristic schedulers inadequate for dynamic resource demands. To address this, we propose SchedTwin—the first real-time digital twin system for HPC job scheduling. It continuously ingests runtime event streams to drive high-fidelity discrete-event simulation, enabling rapid online evaluation of “what-if” scenarios across multiple scheduling policies and facilitating goal-driven, closed-loop adaptive scheduling. Deeply integrated with the PBS scheduler, SchedTwin achieves low-overhead (sub-10-second decision latency) and high-accuracy online policy optimization. Experimental evaluation in production environments demonstrates that SchedTwin significantly outperforms mainstream static schedulers—overcoming the longstanding dual bottlenecks of adaptability and timeliness inherent in conventional HPC scheduling approaches.

Adaptive scheduling for diverse HPC workloadsDynamic policy selection to meet optimization goalsReal-time digital twin guides scheduling decisions

This work addresses the limitations of Kubernetes’ default scheduler, which often leads to resource fragmentation and suboptimal utilization due to its local decision-making nature, while existing global scheduling approaches struggle with practical deployment in production clusters. The paper proposes OPSche, the first open-source plugin that collaboratively operates alongside the default scheduler by leveraging the Kubernetes scheduling framework to seamlessly integrate globally optimized schedules generated by external solvers through atomic validation and coordination hooks. OPSche supports three trigger modes—scheduling failure, periodic invocation, and queue stabilization—along with their blocking variants, thereby balancing scheduling quality, latency, and interference without replacing the native scheduler. Experimental results demonstrate that OPSche improves resource utilization by up to 3.0% across diverse cluster configurations and reduces scheduling latency by over one second.

cluster-wide placementglobal optimizationKubernetes scheduling

Hot Scholars

YY

Yang You

Postdoc, Stanford University
3D visioncomputer graphicscomputational geometry
HS

Haiying Shen

Associate Professor of Computer Science, University of Virginia
Distributed systems and networksDistributed machine learningCloudCPS
MX

Minxian Xu

Associate Professor, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
Cloud ComputingMicroservicesLLM Inference
DL

Dahua Lin

The Chinese University of Hong Kong
computer visionmachine learningprobabilistic inferencebayesian nonparametrics
FF

Fangcheng Fu

Shanghai Jiao Tong University
machine learningdeep learningMLSysdistributed computation