Score
Designs, builds, configures, and operates compute cluster systems and the accompanying automation and tooling, covering cluster architecture, automated provisioning and deployment, lifecycle and resource management, scheduler configuration and scheduling policies (e.g., SLURM), performance tuning, debugging, and multi-cluster federation. Analyzes resource allocation, utilization, job scheduling, and cost metrics to optimize throughput, reliability, and operational cost.
This paper addresses the lack of clarity regarding the diversity and evolutionary trajectories of modern workload schedulers. We propose a cross-layer taxonomy comprising three categories: OS process scheduling, cluster job scheduling, and big-data scheduling. Through algorithmic feature analysis and historical comparative study, we systematically characterize the design rationales, optimization objectives, and technological evolution of these schedulers, uncovering shared design patterns across local and distributed environments. Our key contribution is the first unified classification framework, which identifies three fundamental differentiating dimensions: resource abstraction granularity, scheduling timing, and feedback mechanism. Based on this analysis, we distill general-purpose scheduling design principles targeting heterogeneity, scalability, and QoS guarantees. The study provides both theoretical foundations and practical guidance for scheduler selection, cross-layer coordination optimization, and next-generation scheduler architecture design.
Scientific workflows on clusters suffer from inefficient resource scheduling, high energy consumption, and unpredictable costs due to inaccurate manual performance estimation. To address this, we propose an automated, task-level performance prediction method—estimating both execution time and memory consumption—by integrating machine learning (regression and ensemble models), fine-grained feature engineering, runtime performance modeling, and workflow semantic analysis. Our approach enables cross-platform, multi-objective (including carbon-aware) generalization. We present the first systematic survey and horizontal evaluation of mainstream prediction paradigms, identifying key limitations in dynamism, transferability, and multi-objective coordination, while charting their evolutionary trajectory. We establish a unified benchmarking framework and validate our method on real-world workflows (e.g., CyberShake, SIPHT), achieving 32–47% lower prediction error. This enables resource managers to perform precise scheduling, energy-efficient operation, carbon-aware optimization, and accurate cost estimation—thereby improving cluster resource utilization and scheduling efficiency.
This paper investigates the joint optimization of server count, scheduling policy, and system architecture under a fixed computational budget to minimize average job response time. Using high-resolution traces from Google Cloud production workloads, we develop a multi-stage server cluster model and systematically compare classical policies—including Join-Idle-Queue (JIQ) and Round-Robin (RR)—against state-of-the-art size-aware schedulers. Our findings reveal: (1) an optimal critical server scale that minimizes response time; (2) in high-parallelism or multi-tier architectures, RR and JIQ significantly outperform conventional size-aware policies; and (3) parallelism degree and architectural design exert greater influence on performance than scheduling algorithm sophistication. Collectively, these results establish a new optimization paradigm wherein “architecture–parallelism” dominates over “algorithmic refinement.”
SLURM logs in HPC scientific workflows lack explicit case identifiers, hindering direct application of process mining. Method: This paper proposes an automatic job-correlation method based on implicit job dependency modeling—parsing SLURM logs and jointly leveraging spatiotemporal job feature matching and graph-structured modeling to achieve end-to-end clustering of unannotated jobs. Contribution/Results: We introduce the first systematic preprocessing framework for process mining on HPC logs, integrating algorithms such as Heuristics Miner to support process discovery and bottleneck diagnosis. Evaluated on real-world HPC cluster logs, our approach significantly improves workflow traceability, accurately identifies I/O- and scheduler-related performance bottlenecks, and enables high-fidelity reconstruction of end-to-end process models.
This work addresses the lack of existing tools capable of continuous performance validation and regression detection across entire datacenter clusters. The authors propose the first cluster-wide continuous benchmarking framework that supports unified scheduling, enabling simultaneous task distribution to all nodes and systematic collection of multidimensional performance metrics—spanning CPU, GPU, memory, interconnects, I/O, power consumption, frequency, and temperature—across both space and time. This framework facilitates performance regression detection under software and hardware changes as well as analysis of hardware variability. Experiments on the NHR@FAU cluster reveal intra-node performance variations below 1% among identically configured nodes, while inter-node differences reach up to 5%. The study further uncovers, for the first time, significant disparities in the performance–power relationship between air-cooled and liquid-cooled nodes.
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.
This work addresses the limitations of existing cloud scheduling approaches, which often oversimplify application resource demands and hardware contention, leading to poor resource utilization and performance instability under shared-resource congestion. To overcome these challenges, the authors propose a scheduling framework that integrates a flexible SLO (Service Level Objective) mechanism with hardware-level resource awareness. By permitting brief, controlled SLO violations to avoid over-provisioning and continuously monitoring last-level cache and memory bandwidth congestion, the framework enables informed, resource-aware scheduling and rescheduling decisions. Experimental results demonstrate that, compared to conventional hard-SLO methods, the proposed approach reduces corrective rescheduling by 49% and decreases node-level resource contention by 8%, significantly enhancing cluster-wide efficiency while maintaining performance stability.
This study addresses the challenge of enhancing productivity in supercomputing clusters and informing the design of exascale systems by analyzing job scheduling logs, GPU trace data, and domain-specific metadata from the Titan supercomputer. It systematically investigates the relationship between requested and actual resource utilization and its temporal evolution. Employing correlation analysis, clustering, and neural networks, the work presents the first comprehensive characterization of seasonal patterns in HPC resource usage and develops a transferable model for predicting resource utilization. By identifying key user behavior patterns, the research substantially improves the accuracy of forecasting future resource demands, thereby providing empirical foundations for optimizing configuration and planning of high-performance computing systems.