design runtime schedulers

Designs and implements runtime schedulers, scheduling algorithms, and scheduler implementations that decide task placement, task colocation, degrees of parallelism, deadlines and priorities, and the allocation of compute resources, queues and nodes. Builds and analyzes dynamic, topology-aware, workflow/job/dataflow scheduling policies that perform fine-grained dynamic scheduling, adapt to changing inputs and sparsity patterns, mitigate contention and stalls, orchestrate distributed executions, record run metadata for reproducibility, and optimize schedules for performance, reuse, and fairness.

designruntimeschedulers

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.54
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$205K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Workload Schedulers -- Genesis, Algorithms and Differences

Nov 13, 2025
LS
L. Sliwko
🏛️ University of Westminster

This paper addresses the lack of clarity regarding the diversity and evolutionary trajectories of modern workload schedulers. We propose a cross-layer taxonomy comprising three categories: OS process scheduling, cluster job scheduling, and big-data scheduling. Through algorithmic feature analysis and historical comparative study, we systematically characterize the design rationales, optimization objectives, and technological evolution of these schedulers, uncovering shared design patterns across local and distributed environments. Our key contribution is the first unified classification framework, which identifies three fundamental differentiating dimensions: resource abstraction granularity, scheduling timing, and feedback mechanism. Based on this analysis, we distill general-purpose scheduling design principles targeting heterogeneity, scalability, and QoS guarantees. The study provides both theoretical foundations and practical guidance for scheduler selection, cross-layer coordination optimization, and next-generation scheduler architecture design.

Analyzing scheduler evolution from early adoptions to modern implementationsCategorizing modern workload schedulers into three distinct classesComparing scheduling strategies across local and distributed systems

Memory-aware Adaptive Scheduling of Scientific Workflows on Heterogeneous Architectures

Mar 28, 2025
SK
S. Kulagina
🏛️ Humboldt-Universitaet zu Berlin | ENS Lyon | IUF

To address DAG execution failures and excessive makespan in scientific workflows on heterogeneous platforms caused by runtime memory overflow, this paper proposes a memory-aware, dynamically adaptive scheduling framework. Methodologically: (1) it designs a memory-constrained variant of HEFT integrated with a lightweight data eviction strategy to prevent out-of-memory conditions during execution; (2) it introduces an online rescheduling mechanism triggered by real-time detection of performance parameter deviations, thereby overcoming the limitations of conventional static scheduling assumptions; and (3) it implements an end-to-end, integrated runtime system. Experimental results demonstrate that the framework achieves, for the first time, successful scheduling of previously infeasible (memory-exceeding) workflows, reduces average makespan by 23.6%, and significantly outperforms state-of-the-art schedulers—including HEFT, PEFT, and CPOP—in both feasibility and efficiency.

Adaptive scheduling for dynamic execution time and memory usage changesMinimizing workflow makespan while preventing memory failuresScheduling DAG workflows on heterogeneous platforms with memory constraints

The Merit of Simple Policies: Buying Performance With Parallelism and System Architecture

Mar 20, 2025
MY
Mert Yildiz
🏛️ University of Rome Sapienza

This paper investigates the joint optimization of server count, scheduling policy, and system architecture under a fixed computational budget to minimize average job response time. Using high-resolution traces from Google Cloud production workloads, we develop a multi-stage server cluster model and systematically compare classical policies—including Join-Idle-Queue (JIQ) and Round-Robin (RR)—against state-of-the-art size-aware schedulers. Our findings reveal: (1) an optimal critical server scale that minimizes response time; (2) in high-parallelism or multi-tier architectures, RR and JIQ significantly outperform conventional size-aware policies; and (3) parallelism degree and architectural design exert greater influence on performance than scheduling algorithm sophistication. Collectively, these results establish a new optimization paradigm wherein “architecture–parallelism” dominates over “algorithmic refinement.”

Comparing simple vs. complex dispatching policies for workload scheduling.Exploring the impact of parallelism and system architecture on performance.Optimizing job response time in cloud computing clusters.

This work addresses the challenge of mapping and scheduling large-scale scientific workflows in heterogeneous HPC environments under dynamic constraints and data-awareness requirements. We propose an end-to-end, constraint-aware scheduling framework that tightly integrates Graph Neural Networks (GNNs) with Proximal Policy Optimization (PPO). Workflows are modeled as directed acyclic graphs (DAGs), where GNNs jointly encode task dependencies, resource heterogeneity, and hard constraints (e.g., deadlines, data locality), while PPO enables adaptive, search-free online scheduling via reinforcement learning. Evaluated across multiple real-world datasets, our approach achieves 76% speedup over ILP-based solvers while maintaining near-optimal solution quality, and incurs only a 3.85× runtime overhead compared to the OLB heuristic. The framework seamlessly integrates with SLURM and Kubernetes, strictly satisfies all hard constraints, and demonstrates strong generalization across diverse workflow and infrastructure configurations—effectively balancing optimality and real-time responsiveness.

Handling dynamic constraints and resource requirements efficientlyOptimizing workflow mapping in heterogeneous HPC systemsReducing makespan while maintaining optimal scheduling solutions

WOW: Workflow-Aware Data Movement and Task Scheduling for Dynamic Scientific Workflows

Mar 17, 2025
FL
Fabian Lehmann
🏛️ Humboldt-Universitaet zu Berlin | Technische Universitaet Berlin | Technische Universitaet Darmstadt | University of Glasgow

In dynamic scientific workflows, the decoupling of task scheduling from data movement often assigns tasks to nodes lacking local input data, causing network congestion and execution delays. To address this, we propose the first workflow-aware joint scheduling framework that enables “data-readiness-driven” task placement via predictive pre-staging of intermediate data replicas, supporting dynamic execution plans. Our approach integrates speculative data pre-staging, dependency-aware scheduling, and lightweight storage management, implemented on a Nextflow+Kubernetes prototype. Experiments across 16 synthetic and real-world workflows demonstrate significant reductions in total completion time—up to 94.5% (synthetic) and 53.2% (real), with only bounded, transient storage overhead. The core contribution is the first holistic co-optimization of scheduling and data movement for dynamic workflows, breaking the traditional separation between these concerns.

Improves resource allocation and scalability in cluster environmentsOptimizes task scheduling and data movement in scientific workflowsReduces network congestion and overall workflow runtime

Latest Papers

What's happening recently
View more

This study investigates the scalability and performance of process and thread schedulers under memory-intensive workloads in multi-core shared-memory systems, focusing on a 3D tensor row-sorting task. The authors design and evaluate several scheduling strategies: on the thread side, an AIMD-based adaptive chunking mechanism inspired by TCP congestion control is introduced, coupled with exponential weighted moving average to dynamically adjust concurrency; on the process side, a bounded prolific/collective model is employed alongside one-to-one, one-to-many, and many-to-many pipelined communication patterns to enable flexible task distribution. Experimental results on a 24-core x86-64 platform demonstrate that thread-level scheduling consistently outperforms process-level scheduling, with dynamic and guided strategies achieving the best performance, while the many-to-many pipeline exhibits superior scalability for large-scale tasks.

many-core systemsprocess-based schedulingscalability

This work addresses the limitation of existing distributed data pipeline systems, which require users to explicitly define complete workflow graphs, by proposing a unified planning and scheduling framework that automatically constructs end-to-end persistent pipelines from implicit goal declarations alone. The approach introduces, for the first time, a numeric-domain-independent planner into the context of persistent scheduling, integrating workflow and resource graph modeling, numeric planning, and network interface scheduling to achieve full automation. Experimental results demonstrate the feasibility and scalability of the method: under a single-machine constraint of one hour of CPU time and 30 GB of memory, the system successfully scheduled a linear pipeline spanning eight sites and comprising fourteen components.

automated planningdata pipelinesdistributed workflows

This work addresses the performance degradation of mixed multi-runtime and multi-process workloads under over-subscription, where traditional OS schedulers induce thread interference through periodic preemption, exacerbating lock contention and scalability collapse. To overcome this, the authors propose USF, a user-space scheduling framework that enables cross-process and multi-runtime cooperative scheduling without requiring privileged operations or application modifications. USF employs a cooperative policy, SCHED_COOP, which triggers context switches only when threads voluntarily block, thereby eliminating preemption-induced overheads. Built upon an extended GNU C library and the nOS-V runtime, USF maintains compatibility with mainstream parallel frameworks such as OpenMP. Evaluations on representative workloads—including nested BLAS, multi-process PyTorch with LLaMA-3 inference, and molecular dynamics simulations—demonstrate performance improvements of up to 2.4×.

multi-runtime workloadsOS scheduler interferenceoversubscription

This work addresses the challenge that task scheduling under strong scaling is often constrained by task granularity, where scheduling overhead can dominate performance as parallelism increases, yet a systematic understanding of how algorithmic dependency structures affect scheduling efficiency remains lacking. The paper proposes a novel framework that characterizes granularity based on the dependency topology of task graphs, attributing the growth of scheduling overhead to dependency structure rather than problem size—a distinction made for the first time. Building on this insight, the authors develop a predictive model for strong-scaling limits and derive rules for selecting appropriate scheduling strategies. Through task graph analysis and overhead modeling, the approach accurately explains both gradual and abrupt scaling breakdowns observed across diverse parallel workloads, enabling informed automatic selection between static and dynamic scheduling without exhaustive empirical testing.

dependency topologydynamic schedulingscheduling overhead

Hot Scholars

MG

Minyi Guo

IEEE Fellow, Chair Professor, Shanghai Jiao Tong University
Parallel ComputingCompiler OptimizationCloud ComputingNetworking
XM

Xupeng Miao

Purdue University
Machine Learning SystemsData Management
FF

Fangcheng Fu

Shanghai Jiao Tong University
machine learningdeep learningMLSysdistributed computation
AL

Alexander Lindermayr

Postdoc, Simons Institute, UC Berkeley
algorithmscombinatorial optimizationscheduling
JS

Jens Schlöter

Postdoctoral Researcher, CWI, Amsterdam
Combinatorial optimizationoptimization under uncertaintyapproximation algorithms