Score
Designs, installs, configures, operates, and upgrades cluster job-scheduling systems—primarily Slurm—by managing scheduler daemons, plugins, queue and partition policies, resource limits, accounting, access controls, and node provisioning for multi-node compute clusters. Administers and tunes scheduling parameters, deployment automation, monitoring, and troubleshooting, and analyzes scheduler logs and metrics to diagnose bottlenecks, improve job throughput, utilization, fair-share, and overall scheduler reliability.
This paper addresses the lack of clarity regarding the diversity and evolutionary trajectories of modern workload schedulers. We propose a cross-layer taxonomy comprising three categories: OS process scheduling, cluster job scheduling, and big-data scheduling. Through algorithmic feature analysis and historical comparative study, we systematically characterize the design rationales, optimization objectives, and technological evolution of these schedulers, uncovering shared design patterns across local and distributed environments. Our key contribution is the first unified classification framework, which identifies three fundamental differentiating dimensions: resource abstraction granularity, scheduling timing, and feedback mechanism. Based on this analysis, we distill general-purpose scheduling design principles targeting heterogeneity, scalability, and QoS guarantees. The study provides both theoretical foundations and practical guidance for scheduler selection, cross-layer coordination optimization, and next-generation scheduler architecture design.
This work addresses the high cognitive burden on users and excessive carbon emissions associated with scientific computing due to the complexity of the SLURM job scheduler interface and its lack of energy-aware scheduling mechanisms. To mitigate these issues, the authors propose a modular Perl-based toolkit featuring a simplified command-line interface and a text-based user interface (TUI) that supports job monitoring, cancellation, and automatic generation of specialized submission scripts. A key innovation is the introduction of an “eco-mode” that enables automatic energy-efficient scheduling through off-peak workload shifting. This approach significantly lowers the usability barrier, enhances job management efficiency, and effectively reduces the carbon footprint of research computing workflows.
This study addresses scheduler RPC overload in shared Slurm clusters caused by large-scale Nextflow workloads. We propose a reproducible measurement protocol and unified timing attribution method to quantitatively evaluate performance trade-offs across four deployment strategies. By benchmarking Nextflow, HyperQueue, and Flux, we characterize the Pareto frontier between wall-clock time and RPC overhead. Our results demonstrate that Flux significantly reduces scheduling load, establishing it as an optimal solution for high-performance computing sites. These findings provide quantitative guidance for selecting deployment strategies that maintain scheduler stability while preserving user experience in multi-tenant HPC environments.
为解决HPC集群上执行自主工作流的问题,本文提出了RASER框架,通过扩展Slurm内部机制实现动态任务调度和容错。
This work addresses the limitations of existing cloud scheduling approaches, which often oversimplify application resource demands and hardware contention, leading to poor resource utilization and performance instability under shared-resource congestion. To overcome these challenges, the authors propose a scheduling framework that integrates a flexible SLO (Service Level Objective) mechanism with hardware-level resource awareness. By permitting brief, controlled SLO violations to avoid over-provisioning and continuously monitoring last-level cache and memory bandwidth congestion, the framework enables informed, resource-aware scheduling and rescheduling decisions. Experimental results demonstrate that, compared to conventional hard-SLO methods, the proposed approach reduces corrective rescheduling by 49% and decreases node-level resource contention by 8%, significantly enhancing cluster-wide efficiency while maintaining performance stability.
This work addresses the inefficiencies of traditional rigid job scheduling in high-performance computing (HPC) clusters, which often result in low resource utilization and prolonged job waiting times. The authors propose a novel malleable job scheduling strategy that dynamically adjusts resource allocations at runtime while prioritizing each job’s preferred configuration. They systematically investigate the interplay among workload characteristics, the proportion of malleable jobs, and scheduling policies. Using the ElastiSim simulation framework and real-world workload traces from the Cori, Eagle, and Theta supercomputers, they evaluate five scheduling strategies across malleable job ratios ranging from 0% to 100%. Experimental results demonstrate that, compared to fully rigid scheduling, the best-performing strategy reduces job turnaround time by 37–67%, shortens makespan by 16–65%, decreases waiting time by 73–99%, and improves node utilization by 5–52%.
This work addresses the limitations of Kubernetes’ default scheduler, which often leads to resource fragmentation and suboptimal utilization due to its local decision-making nature, while existing global scheduling approaches struggle with practical deployment in production clusters. The paper proposes OPSche, the first open-source plugin that collaboratively operates alongside the default scheduler by leveraging the Kubernetes scheduling framework to seamlessly integrate globally optimized schedules generated by external solvers through atomic validation and coordination hooks. OPSche supports three trigger modes—scheduling failure, periodic invocation, and queue stabilization—along with their blocking variants, thereby balancing scheduling quality, latency, and interference without replacing the native scheduler. Experimental results demonstrate that OPSche improves resource utilization by up to 3.0% across diverse cluster configurations and reduces scheduling latency by over one second.
This study addresses the challenges of service coordination and data auditing in message-passing experiments on high-performance computing (HPC) clusters by proposing a Slurm-orchestrated benchmarking framework. Methodologically, it introduces an "eligibility-first" evaluation paradigm that enforces data validity verification prior to throughput assessment, ensuring results satisfy causal logic constraints. Technically, the framework integrates a single-broker Kafka architecture, in-memory log storage, and multi-stage repeated validation mechanisms to achieve recoverable experiment control and distributed auditing. Experimental evaluations across 120 workloads yield 99 eligible observations, with selected workloads achieving a 100% qualification rate. Furthermore, the system attains a balanced endpoint throughput of 2,927 MiB/s with a P99 latency of 2.56 seconds, demonstrating both robust auditability and high performance.
This work addresses the lack of a secure, efficient, and compatible programmatic interface in existing Slurm schedulers, which hinders integration with scientific tools and automation systems. We propose Palmetto API, a lightweight RESTful proxy for slurmrestd that introduces, for the first time, fine-grained role-based access control (RBAC) and HTTP response caching while maintaining full compatibility with existing clients. This design significantly enhances interface security and performance, effectively reducing latency from redundant requests. Compatibility validation confirms seamless interoperability with the current Slurm ecosystem, enabling straightforward adoption without disrupting established workflows.
This work addresses the technical and behavioral challenges of transitioning from node-exclusive to resource-aware scheduling in production-grade heterogeneous HPC systems, a shift that risks disrupting established scientific workflows. To enable seamless, non-disruptive migration, the authors propose a collaborative operational framework integrating a time-bound compatibility layer, observability-driven feedback mechanisms, and targeted user guidance. Built upon Slurm’s TRES resource model, the approach combines runtime compatibility support, job queue monitoring, and user behavior analysis to preserve workflow continuity while substantially improving scheduling efficiency. Empirical results demonstrate dramatic reductions in median queue wait times—from 277 minutes to under 3 minutes for CPU jobs and from 81 minutes to 3.4 minutes for GPU jobs—alongside high long-term adoption rates among users who embraced the new submission paradigm.