cluster scheduler administration

Designs, installs, configures, operates, and upgrades cluster job-scheduling systems—primarily Slurm—by managing scheduler daemons, plugins, queue and partition policies, resource limits, accounting, access controls, and node provisioning for multi-node compute clusters. Administers and tunes scheduling parameters, deployment automation, monitoring, and troubleshooting, and analyzes scheduler logs and metrics to diagnose bottlenecks, improve job throughput, utilization, fair-share, and overall scheduler reliability.

clusterscheduleradministration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Workload Schedulers -- Genesis, Algorithms and Differences

Nov 13, 2025
LS
L. Sliwko
🏛️ University of Westminster

This paper addresses the lack of clarity regarding the diversity and evolutionary trajectories of modern workload schedulers. We propose a cross-layer taxonomy comprising three categories: OS process scheduling, cluster job scheduling, and big-data scheduling. Through algorithmic feature analysis and historical comparative study, we systematically characterize the design rationales, optimization objectives, and technological evolution of these schedulers, uncovering shared design patterns across local and distributed environments. Our key contribution is the first unified classification framework, which identifies three fundamental differentiating dimensions: resource abstraction granularity, scheduling timing, and feedback mechanism. Based on this analysis, we distill general-purpose scheduling design principles targeting heterogeneity, scalability, and QoS guarantees. The study provides both theoretical foundations and practical guidance for scheduler selection, cross-layer coordination optimization, and next-generation scheduler architecture design.

Analyzing scheduler evolution from early adoptions to modern implementationsCategorizing modern workload schedulers into three distinct classesComparing scheduling strategies across local and distributed systems

This work addresses the high cognitive burden on users and excessive carbon emissions associated with scientific computing due to the complexity of the SLURM job scheduler interface and its lack of energy-aware scheduling mechanisms. To mitigate these issues, the authors propose a modular Perl-based toolkit featuring a simplified command-line interface and a text-based user interface (TUI) that supports job monitoring, cancellation, and automatic generation of specialized submission scripts. A key innovation is the introduction of an “eco-mode” that enables automatic energy-efficient scheduling through off-peak workload shifting. This approach significantly lowers the usability barrier, enhances job management efficiency, and effectively reduces the carbon footprint of research computing workflows.

carbon footprintenergy savingHPC

This study addresses scheduler RPC overload in shared Slurm clusters caused by large-scale Nextflow workloads. We propose a reproducible measurement protocol and unified timing attribution method to quantitatively evaluate performance trade-offs across four deployment strategies. By benchmarking Nextflow, HyperQueue, and Flux, we characterize the Pareto frontier between wall-clock time and RPC overhead. Our results demonstrate that Flux significantly reduces scheduling load, establishing it as an optimal solution for high-performance computing sites. These findings provide quantitative guidance for selecting deployment strategies that maintain scheduler stability while preserving user experience in multi-tenant HPC environments.

deployment strategiesNextflowRPC overhead

This work addresses the limitations of existing cloud scheduling approaches, which often oversimplify application resource demands and hardware contention, leading to poor resource utilization and performance instability under shared-resource congestion. To overcome these challenges, the authors propose a scheduling framework that integrates a flexible SLO (Service Level Objective) mechanism with hardware-level resource awareness. By permitting brief, controlled SLO violations to avoid over-provisioning and continuously monitoring last-level cache and memory bandwidth congestion, the framework enables informed, resource-aware scheduling and rescheduling decisions. Experimental results demonstrate that, compared to conventional hard-SLO methods, the proposed approach reduces corrective rescheduling by 49% and decreases node-level resource contention by 8%, significantly enhancing cluster-wide efficiency while maintaining performance stability.

cloud computingcluster schedulingresource congestion

Latest Papers

What's happening recently
View more

This work addresses the inefficiencies of traditional rigid job scheduling in high-performance computing (HPC) clusters, which often result in low resource utilization and prolonged job waiting times. The authors propose a novel malleable job scheduling strategy that dynamically adjusts resource allocations at runtime while prioritizing each job’s preferred configuration. They systematically investigate the interplay among workload characteristics, the proportion of malleable jobs, and scheduling policies. Using the ElastiSim simulation framework and real-world workload traces from the Cori, Eagle, and Theta supercomputers, they evaluate five scheduling strategies across malleable job ratios ranging from 0% to 100%. Experimental results demonstrate that, compared to fully rigid scheduling, the best-performing strategy reduces job turnaround time by 37–67%, shortens makespan by 16–65%, decreases waiting time by 73–99%, and improves node utilization by 5–52%.

HPC clustersjob waiting timemalleable job scheduling

This work addresses the limitations of Kubernetes’ default scheduler, which often leads to resource fragmentation and suboptimal utilization due to its local decision-making nature, while existing global scheduling approaches struggle with practical deployment in production clusters. The paper proposes OPSche, the first open-source plugin that collaboratively operates alongside the default scheduler by leveraging the Kubernetes scheduling framework to seamlessly integrate globally optimized schedules generated by external solvers through atomic validation and coordination hooks. OPSche supports three trigger modes—scheduling failure, periodic invocation, and queue stabilization—along with their blocking variants, thereby balancing scheduling quality, latency, and interference without replacing the native scheduler. Experimental results demonstrate that OPSche improves resource utilization by up to 3.0% across diverse cluster configurations and reduces scheduling latency by over one second.

cluster-wide placementglobal optimizationKubernetes scheduling

This study addresses the challenges of service coordination and data auditing in message-passing experiments on high-performance computing (HPC) clusters by proposing a Slurm-orchestrated benchmarking framework. Methodologically, it introduces an "eligibility-first" evaluation paradigm that enforces data validity verification prior to throughput assessment, ensuring results satisfy causal logic constraints. Technically, the framework integrates a single-broker Kafka architecture, in-memory log storage, and multi-stage repeated validation mechanisms to achieve recoverable experiment control and distributed auditing. Experimental evaluations across 120 workloads yield 99 eligible observations, with selected workloads achieving a 100% qualification rate. Furthermore, the system attains a balanced endpoint throughput of 2,927 MiB/s with a P99 latency of 2.56 seconds, demonstrating both robust auditability and high performance.

Experiment QualificationHigh-Performance ComputingKafka Evaluation

This work addresses the lack of a secure, efficient, and compatible programmatic interface in existing Slurm schedulers, which hinders integration with scientific tools and automation systems. We propose Palmetto API, a lightweight RESTful proxy for slurmrestd that introduces, for the first time, fine-grained role-based access control (RBAC) and HTTP response caching while maintaining full compatibility with existing clients. This design significantly enhances interface security and performance, effectively reducing latency from redundant requests. Compatibility validation confirms seamless interoperability with the current Slurm ecosystem, enabling straightforward adoption without disrupting established workflows.

cachingcluster schedulercompatibility

This work addresses the technical and behavioral challenges of transitioning from node-exclusive to resource-aware scheduling in production-grade heterogeneous HPC systems, a shift that risks disrupting established scientific workflows. To enable seamless, non-disruptive migration, the authors propose a collaborative operational framework integrating a time-bound compatibility layer, observability-driven feedback mechanisms, and targeted user guidance. Built upon Slurm’s TRES resource model, the approach combines runtime compatibility support, job queue monitoring, and user behavior analysis to preserve workflow continuity while substantially improving scheduling efficiency. Empirical results demonstrate dramatic reductions in median queue wait times—from 277 minutes to under 3 minutes for CPU jobs and from 81 minutes to 3.4 minutes for GPU jobs—alongside high long-term adoption rates among users who embraced the new submission paradigm.

non-disruptive migrationproduction HPCresource-aware scheduling

Hot Scholars

MP

Marco Pegoraro

PhD student at Sapienza University of Rome
Deep learningGeometry ProcessingStructural Biology