gang scheduling

Designs, implements, and evaluates scheduling algorithms and runtime mechanisms that allocate and coordinate groups of related tasks or processes to run simultaneously across multiple CPUs or nodes. Works on gang launch, synchronization, preemption, migration and admission policies and measures their impact on utilization, synchronization wait, job completion time, fairness, overhead and robustness under contention and failures.

gangscheduling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Workload Schedulers -- Genesis, Algorithms and Differences

Nov 13, 2025
LS
L. Sliwko
🏛️ University of Westminster

This paper addresses the lack of clarity regarding the diversity and evolutionary trajectories of modern workload schedulers. We propose a cross-layer taxonomy comprising three categories: OS process scheduling, cluster job scheduling, and big-data scheduling. Through algorithmic feature analysis and historical comparative study, we systematically characterize the design rationales, optimization objectives, and technological evolution of these schedulers, uncovering shared design patterns across local and distributed environments. Our key contribution is the first unified classification framework, which identifies three fundamental differentiating dimensions: resource abstraction granularity, scheduling timing, and feedback mechanism. Based on this analysis, we distill general-purpose scheduling design principles targeting heterogeneity, scalability, and QoS guarantees. The study provides both theoretical foundations and practical guidance for scheduler selection, cross-layer coordination optimization, and next-generation scheduler architecture design.

Analyzing scheduler evolution from early adoptions to modern implementationsCategorizing modern workload schedulers into three distinct classesComparing scheduling strategies across local and distributed systems

This work addresses the challenge of balancing performance and fairness in dynamic task graph scheduling, where traditional approaches often neglect adjustments to existing task assignments. To overcome this limitation, the authors propose the Last-K Preemption model, which introduces a controlled, localized preemption mechanism that reschedules only the most recent K task graphs while preserving earlier allocations. This strategy effectively balances scheduling efficiency against system overhead. Extensive experiments are conducted using synthetic, RIoTBench, WFCommons, and adversarial workloads, comparing fully preemptive, non-preemptive, and partially preemptive strategies. The results demonstrate that the proposed moderate preemption approach achieves makespan and resource utilization comparable to full preemption, while significantly reducing scheduling overhead and ensuring fairness.

dynamic task graph schedulingfairnessmakespan

Minimize Your Critical Path with Combine-and-Exchange Locks

Nov 12, 2025
SK
Simon König
🏛️ University of Stuttgart

Existing user-space coroutine/fiber synchronization mechanisms implicitly assume kernel scheduling, introducing unnecessary latency on critical paths and limiting high-concurrency throughput. This paper proposes Combine-and-Exchange Scheduling (CES), a novel synchronization paradigm for purely user-space cooperative scheduling. CES eliminates cross-thread overhead by retaining critical sections on the same thread during lock contention, while dynamically redistributing parallelizable tasks to idle threads. Crucially, it co-designs user-space synchronization primitives with the scheduler to fully bypass kernel intervention. Experimental evaluation demonstrates that CES achieves up to 3× higher throughput on application-level benchmarks and up to 8× speedup on microbenchmarks—significantly outperforming state-of-the-art user-space synchronization approaches.

Improving throughput for coroutine-based parallel applicationsOptimizing scheduling for contended critical sections across threadsReducing critical path delays in userspace synchronization primitives

This study investigates the scalability and performance of process and thread schedulers under memory-intensive workloads in multi-core shared-memory systems, focusing on a 3D tensor row-sorting task. The authors design and evaluate several scheduling strategies: on the thread side, an AIMD-based adaptive chunking mechanism inspired by TCP congestion control is introduced, coupled with exponential weighted moving average to dynamically adjust concurrency; on the process side, a bounded prolific/collective model is employed alongside one-to-one, one-to-many, and many-to-many pipelined communication patterns to enable flexible task distribution. Experimental results on a 24-core x86-64 platform demonstrate that thread-level scheduling consistently outperforms process-level scheduling, with dynamic and guided strategies achieving the best performance, while the many-to-many pipeline exhibits superior scalability for large-scale tasks.

many-core systemsprocess-based schedulingscalability

This paper studies the scheduling of multi-class parallelizable jobs under limited server resources to minimize average response time. Jobs are categorized by parallelizability and size distribution, and must be dynamically assigned to $k$ servers. The work first reveals that the load regime fundamentally determines the optimal policy: Least-Parallelizable-First is asymptotically optimal in the sub-Halfin–Whitt (light-load) regime, while SERPT is asymptotically optimal in the super-NDS (heavy-load) regime. Leveraging this insight, we design an adaptive policy that requires no prior knowledge of system load. Using tools from multi-class queueing theory, load-scaling analysis, and online algorithm design, we rigorously prove that each regime-specific policy achieves the fundamental lower bound on average response time. Simulation results demonstrate that the proposed adaptive strategy consistently attains near-theoretically-optimal performance across the entire load spectrum.

Determines optimal scheduling policies based on job parallelizability and load conditionsMinimizes mean response time under varying load regimesOptimizes server allocation for multiple parallelizable job classes

Latest Papers

What's happening recently
View more

This work addresses the performance degradation of mixed multi-runtime and multi-process workloads under over-subscription, where traditional OS schedulers induce thread interference through periodic preemption, exacerbating lock contention and scalability collapse. To overcome this, the authors propose USF, a user-space scheduling framework that enables cross-process and multi-runtime cooperative scheduling without requiring privileged operations or application modifications. USF employs a cooperative policy, SCHED_COOP, which triggers context switches only when threads voluntarily block, thereby eliminating preemption-induced overheads. Built upon an extended GNU C library and the nOS-V runtime, USF maintains compatibility with mainstream parallel frameworks such as OpenMP. Evaluations on representative workloads—including nested BLAS, multi-process PyTorch with LLaMA-3 inference, and molecular dynamics simulations—demonstrate performance improvements of up to 2.4×.

multi-runtime workloadsOS scheduler interferenceoversubscription

This study addresses the trade-off between computational overload and latency caused by immediate task acceptance in dynamic multi-robot task allocation. We propose a batch-based task acceptance merging mechanism integrated with the Consensus-Based Auction Algorithm (CBAA). Hardware-in-the-loop experiments conducted on AGX Orin and RP2040 platforms systematically quantify the nonlinear effects of merging strategies on latency and computational load under varying processing capacities. Results demonstrate that this mechanism significantly reduces processor load in high-demand scenarios but introduces additional latency when tasks arrive sparsely, revealing the platform-dependent efficacy of merging strategies. These findings provide empirical evidence for real-time task scheduling in resource-constrained systems.

Decentralized systemsMulti-robot task allocationProcessor workload

This study addresses a critical limitation in classical queueing analysis—its frequent neglect of preemption overhead—which hinders accurate assessment of stability and response time in preemptive scheduling systems. Focusing on the M/G/1 queue with preemption overhead, this work investigates class-based preemptive priority scheduling and presents the first exact analysis of response time distributions for such systems. By introducing a novel theoretical construct termed “task joint transform,” which integrates Laplace transforms with stochastic process techniques, the authors derive recursive formulas for the Laplace transforms of response times for tasks of arbitrary classes. This framework enables closed-form computation of all response time moments, clearly elucidates the performance impact of preemption overhead, and establishes a general analytical foundation extendable to broader scheduling overhead models.

M/G/1 queuepreemption overheadpriority scheduling

This work addresses the challenge that task scheduling under strong scaling is often constrained by task granularity, where scheduling overhead can dominate performance as parallelism increases, yet a systematic understanding of how algorithmic dependency structures affect scheduling efficiency remains lacking. The paper proposes a novel framework that characterizes granularity based on the dependency topology of task graphs, attributing the growth of scheduling overhead to dependency structure rather than problem size—a distinction made for the first time. Building on this insight, the authors develop a predictive model for strong-scaling limits and derive rules for selecting appropriate scheduling strategies. Through task graph analysis and overhead modeling, the approach accurately explains both gradual and abrupt scaling breakdowns observed across diverse parallel workloads, enabling informed automatic selection between static and dynamic scheduling without exhaustive empirical testing.

dependency topologydynamic schedulingscheduling overhead

Hot Scholars

DM

Daniel Milroy

Lawrence Livermore National Laboratory
Numerical AnalysisParallel Computing
HT

Hiroyuki Takizawa

Tohoku University
High Performance ComputingSystem SoftwareProgrammingComputer Architecture
RP

Reza Pulungan

Department of Computer Science and Electronics, Universitas Gadjah Mada
Stochastic processesMarkov processesPhase-type distributionsModelling and analysis of concurrent and network
VS

Vanessa Sochat

Lawrence Livermore National Laboratory
software engineeringcontainersschedulingKubernetes