Score
Design, build, and evaluate mechanisms that detect, tolerate, and reduce the impact of slow or delayed workers ("stragglers") in parallel and distributed computations. This includes implementing and measuring approaches such as speculative execution, task replication or coding, scheduling and monitoring policies, and analyzing trade-offs between latency, resource overhead, and correctness.
This paper addresses the lack of clarity regarding the diversity and evolutionary trajectories of modern workload schedulers. We propose a cross-layer taxonomy comprising three categories: OS process scheduling, cluster job scheduling, and big-data scheduling. Through algorithmic feature analysis and historical comparative study, we systematically characterize the design rationales, optimization objectives, and technological evolution of these schedulers, uncovering shared design patterns across local and distributed environments. Our key contribution is the first unified classification framework, which identifies three fundamental differentiating dimensions: resource abstraction granularity, scheduling timing, and feedback mechanism. Based on this analysis, we distill general-purpose scheduling design principles targeting heterogeneity, scalability, and QoS guarantees. The study provides both theoretical foundations and practical guidance for scheduler selection, cross-layer coordination optimization, and next-generation scheduler architecture design.
To address elevated latency in Java servers caused by CPU contention between garbage collection (GC) threads and application threads, this paper introduces an opportunistic scheduling mechanism—first implemented in ZGC—that dynamically schedules GC work exclusively during CPU idle periods, thereby prioritizing application thread execution. The approach preserves ZGC’s core algorithm and requires only lightweight modifications: an extension to the Linux CFS scheduler and minimal changes to ZGC’s source code—ensuring low overhead and high compatibility. Evaluation on SPECjbb2015 shows a 15% throughput improvement under a 25 ms latency bound; Hazelcast benchmarks demonstrate a 40% reduction in mean latency. Across multi-workload scenarios, service responsiveness and throughput stability are significantly enhanced. This work provides a novel, scale-free SLA optimization path for latency-sensitive Java services operating under moderate load.
This work addresses performance degradation and stability deterioration in barrier-mode parallel systems (e.g., Spark’s Barrier Execution Mode) caused by synchronization barriers. We propose a systematic modeling and analytical framework, introducing the first stochastic process model for $(s,k,l)$-redundant barrier systems and deriving rigorous stability conditions and performance upper bounds using queueing theory. We identify dual-event triggering combined with polling-based scheduling as the dominant source of overhead, and establish analytically tractable performance bounds for mixed workloads—both with and without barrier tasks. The model is validated via distribution fitting, overhead attribution analysis, and empirical measurements on Spark; simulation and measurement results show latency distribution errors under 8%. Our contributions provide a theoretical foundation and quantitative toolkit for designing, optimizing, and ensuring stability in barrier-synchronized parallel systems.
This work addresses the challenge of detecting “slow faults”—subtle node degradations that silently impair performance but evade conventional health checks in large-scale model training. The authors propose a novel health management system that synergistically combines lightweight online performance monitoring with offline, systematic node scanning to jointly identify both abrupt failures and long-term slow faults. By leveraging runtime performance tracing, FLOPs utilization analysis, and exhaustive node-level stress testing, the approach substantially enhances system observability and scalability. Experimental results demonstrate that deployment of this method increases peak FLOPs utilization by up to 1.7×, reduces training step duration variance from 20% to 1%, and effectively extends mean time between failures while lowering operational overhead.
This study addresses the problem of scheduling delay-sensitive tasks to spot and on-demand cloud instances under an average latency constraint, aiming to minimize average cost. By modeling the system using queueing theory and stochastic processes, and leveraging convex optimization and knapsack problem analysis, the work characterizes the optimal scheduling structure in both low- and high-latency regimes: it proves that a queue length of one is optimal in the former, while in the latter, it designs an approximation-optimal policy based on knapsack formulation. An adaptive scheduling algorithm is further proposed to dynamically exploit the allowable latency window. Experimental results demonstrate that the algorithm achieves near-theoretical-optimal cost while effectively balancing latency constraints and resource expenditure. This work provides the first analytical solution for scheduling across hybrid spot and on-demand instances under latency guarantees.
This work addresses the heightened risk of silent data corruption (SDC) in large-scale supercomputing clusters, where existing replication-based fault-tolerance mechanisms struggle to accommodate asynchronous many-task (AMT) runtimes that support dynamic task generation and work stealing. The authors propose a lightweight SDC detection and recovery mechanism tailored for nested fork-join programs. By recording the task dependency tree, comparing results, and performing a top-down identification of corrupted tasks, the approach selectively re-executes only the affected tasks while reusing correct results from their subtasks. This method achieves precise, localized recovery under dynamic scheduling for the first time, substantially reducing fault-tolerance overhead. Experimental results demonstrate negligible detection and recovery costs, correctness guarantees, and extensibility to future-based task models.
This work addresses silent data corruption (SDC) in highly dynamic asynchronous multitasking (AMT) environments characterized by runtime task generation and evolving task dependencies. It proposes a tightly coupled redundancy mechanism that integrates primary and replica computations within an asynchronous task runtime based on C++11 futures/promises and work-stealing load balancing. The approach performs runtime cross-validation of all output effects and selectively re-executes only the affected tasks upon SDC detection. As the first solution to enable efficient SDC protection in such dynamic AMT settings, it supports conditional task spawning and dynamic task graphs while incurring less than 2× performance overhead in fault-free execution—partly due to improved load balancing—and achieves SDC recovery with an overhead of approximately 0.5% of total execution time per incident.
This study addresses the problem of burst-induced job wait time exacerbation in cross-facility scientific workflows, leveraging production logs from the Jean Zay supercomputer. Methodologically, it reveals the independence between burst penalties and Slurm’s fair-share factor, proposes a short-cycle throttling mechanism, and constructs a model capturing the strong correlation between job wait times and intra-burst rankings. These findings are quantitatively validated through Pearson correlation analysis, convolution formula derivation, and Welch’s t-tests. Results demonstrate that over 40% of jobs exhibit wait times highly correlated with their burst rankings. Furthermore, partition occupancy and requested core counts are identified as critical influencing factors. This work provides empirical evidence for optimizing scheduling fairness in high-performance computing environments.
This work addresses the challenge of maintaining both timeliness and stability in real-time data streams within scientific workflows, which are highly susceptible to hardware failures, network disruptions, and performance fluctuations in complex environments. The authors propose a lightweight, non-intrusive fault-tolerance mechanism that integrates asynchronous, non-blocking checkpointing with a progress-aware dynamic load redistribution strategy. This approach enables efficient fault recovery and resource rebalancing without interrupting ongoing computations. Under fault-free conditions, the method incurs less than 1% runtime overhead, while in high-failure-rate scenarios, it reduces the impact of faults and performance anomalies by up to sixfold, substantially enhancing the resilience and resource utilization of stream processing systems.
This work addresses the lack of resource-centric computational efficiency metrics—specifically in terms of node-hours—for existing supercomputers and large-scale AI training platforms operating under high failure rates. It proposes the first efficiency evaluation framework grounded in resource consumption rather than execution time, unifying failure rate, mean time between failures, and checkpoint/restart overhead into a cohesive resource-based model. The framework extends Daly’s (2006) model to accommodate heterogeneous scientific workloads. Validated on one year of production data from the Frontier supercomputer, the approach leverages runtime log analysis, joint modeling of failures and checkpointing, and optimization algorithms to accurately quantify the expected fraction of resources usable for scientific computation and to determine optimal checkpoint intervals that minimize resource loss.