Score
Designs and implements cluster-level workload orchestration and scheduling solutions that deploy, configure, and integrate SLURM and Kubernetes to manage job submission, resource allocation, queues, dependencies, and node-level orchestration. Builds scheduling policies, job lifecycle tooling, resource accounting, autoscaling, and connectors (for SLURM↔Kubernetes interoperability) to run batch, parallel, and containerized workloads efficiently across compute clusters.
This study addresses scheduler RPC overload in shared Slurm clusters caused by large-scale Nextflow workloads. We propose a reproducible measurement protocol and unified timing attribution method to quantitatively evaluate performance trade-offs across four deployment strategies. By benchmarking Nextflow, HyperQueue, and Flux, we characterize the Pareto frontier between wall-clock time and RPC overhead. Our results demonstrate that Flux significantly reduces scheduling load, establishing it as an optimal solution for high-performance computing sites. These findings provide quantitative guidance for selecting deployment strategies that maintain scheduler stability while preserving user experience in multi-tenant HPC environments.
This work addresses the high cognitive burden on users and excessive carbon emissions associated with scientific computing due to the complexity of the SLURM job scheduler interface and its lack of energy-aware scheduling mechanisms. To mitigate these issues, the authors propose a modular Perl-based toolkit featuring a simplified command-line interface and a text-based user interface (TUI) that supports job monitoring, cancellation, and automatic generation of specialized submission scripts. A key innovation is the introduction of an “eco-mode” that enables automatic energy-efficient scheduling through off-peak workload shifting. This approach significantly lowers the usability barrier, enhances job management efficiency, and effectively reduces the carbon footprint of research computing workflows.
In high-density Linux clusters, frequent CPU context switches cause significant performance degradation; even with optimal scheduler placement policies, excessive resource over-provisioning is commonly relied upon for mitigation—leading to substantial waste. This paper proposes a latency-aware group scheduling optimization: departing from traditional per-task fairness prioritization, it instead uses task completion latency as the primary scheduling objective. Leveraging dynamic cgroup workload characterization, it adaptively regulates runqueues and deeply modifies the Linux kernel scheduler to enable fine-grained, low-overhead group-level scheduling. Experimental evaluation demonstrates that, while strictly satisfying service-level agreement (SLA) constraints, the approach reduces cluster resource requirements by 28%, markedly improving resource utilization and overall system throughput.
This work addresses the limitations of Kubernetes’ default scheduler, which often leads to resource fragmentation and suboptimal utilization due to its local decision-making nature, while existing global scheduling approaches struggle with practical deployment in production clusters. The paper proposes OPSche, the first open-source plugin that collaboratively operates alongside the default scheduler by leveraging the Kubernetes scheduling framework to seamlessly integrate globally optimized schedules generated by external solvers through atomic validation and coordination hooks. OPSche supports three trigger modes—scheduling failure, periodic invocation, and queue stabilization—along with their blocking variants, thereby balancing scheduling quality, latency, and interference without replacing the native scheduler. Experimental results demonstrate that OPSche improves resource utilization by up to 3.0% across diverse cluster configurations and reduces scheduling latency by over one second.
SLURM logs in HPC scientific workflows lack explicit case identifiers, hindering direct application of process mining. Method: This paper proposes an automatic job-correlation method based on implicit job dependency modeling—parsing SLURM logs and jointly leveraging spatiotemporal job feature matching and graph-structured modeling to achieve end-to-end clustering of unannotated jobs. Contribution/Results: We introduce the first systematic preprocessing framework for process mining on HPC logs, integrating algorithms such as Heuristics Miner to support process discovery and bottleneck diagnosis. Evaluated on real-world HPC cluster logs, our approach significantly improves workflow traceability, accurately identifies I/O- and scheduler-related performance bottlenecks, and enables high-fidelity reconstruction of end-to-end process models.
This work addresses the limitations of existing cloud scheduling approaches, which often oversimplify application resource demands and hardware contention, leading to poor resource utilization and performance instability under shared-resource congestion. To overcome these challenges, the authors propose a scheduling framework that integrates a flexible SLO (Service Level Objective) mechanism with hardware-level resource awareness. By permitting brief, controlled SLO violations to avoid over-provisioning and continuously monitoring last-level cache and memory bandwidth congestion, the framework enables informed, resource-aware scheduling and rescheduling decisions. Experimental results demonstrate that, compared to conventional hard-SLO methods, the proposed approach reduces corrective rescheduling by 49% and decreases node-level resource contention by 8%, significantly enhancing cluster-wide efficiency while maintaining performance stability.
为解决HPC集群上执行自主工作流的问题,本文提出了RASER框架,通过扩展Slurm内部机制实现动态任务调度和容错。
This study addresses the challenges of service coordination and data auditing in message-passing experiments on high-performance computing (HPC) clusters by proposing a Slurm-orchestrated benchmarking framework. Methodologically, it introduces an "eligibility-first" evaluation paradigm that enforces data validity verification prior to throughput assessment, ensuring results satisfy causal logic constraints. Technically, the framework integrates a single-broker Kafka architecture, in-memory log storage, and multi-stage repeated validation mechanisms to achieve recoverable experiment control and distributed auditing. Experimental evaluations across 120 workloads yield 99 eligible observations, with selected workloads achieving a 100% qualification rate. Furthermore, the system attains a balanced endpoint throughput of 2,927 MiB/s with a P99 latency of 2.56 seconds, demonstrating both robust auditability and high performance.
This work addresses the scheduling challenge of simultaneously achieving KV cache reuse, latency SLO compliance, and GPU resource minimization in large-scale LLM serving for agents. We propose an interference-aware request packing mechanism centered on a compact white-box model that accurately predicts interference-induced latency during both prefill and decode phases. Leveraging these predictions, we design an SLO-aware scheduling algorithm that efficiently consolidates requests onto fewer instances while preserving cache reuse, thereby improving per-GPU throughput. Experimental results demonstrate that our approach reduces GPU hours by 16.8%–24.6% compared to state-of-the-art baselines. Furthermore, deployment in a production cluster comprising over one thousand H20 GPUs achieves a 34.7% reduction in resource overhead while strictly satisfying time-per-output-token (TPOT) targets.
This work addresses the inefficiency of existing LLM serving systems, which scale entire models as monolithic units and thus struggle to handle bursty workloads, often leading to SLO violations or underutilized GPU resources. To overcome these limitations, the paper introduces the first operator-level elasticity framework that exploits heterogeneity and elasticity across individual model operators. By co-optimizing operator-level performance profiling, resource provisioning, deployment policies, and runtime scheduling, the proposed approach achieves fine-grained resource management beyond conventional model-level scaling. Experiments on clusters with 40 A100 and 24 GB200 GPUs demonstrate that the system reduces GPU usage by up to 36.3% and power consumption by 28% compared to baselines, or alternatively improves throughput by 44% under fixed hardware costs.