Score
Designs, implements, and evaluates algorithms, heuristics, and runtime mechanisms that monitor per-worker progress and dynamically redistribute or migrate tasks to reduce execution skew and mitigate stragglers. Analyzes load balancing strategies, adaptive redistribution policies, and performance trade-offs to maintain steady throughput and balanced resource utilization under changing workloads.
To address node overload, high operational costs, and poor system stability caused by dynamic heterogeneous resource scheduling in cloud computing, this paper proposes an intelligent load-balancing framework. The method constructs a high-fidelity simulation environment and an abstracted multi-resource model, introducing for the first time a joint resource utilization metric that incorporates VM migration overhead. It establishes a novel three-category taxonomy for schedulers, derives an empirically grounded formula for estimating VM migration traffic, and comparatively evaluates two emerging paradigms: centralized metaheuristic and distributed multi-agent scheduling. Built upon real-world Google cluster traces, the framework integrates live VM migration and realistic workload simulation. Experimental validation on the University of Westminster’s HPC cluster demonstrates a 23.6% improvement in resource utilization, a 31.4% reduction in task latency, and a 27.9% decrease in network migration overhead—significantly enhancing system stability and cost-efficiency.
This paper addresses the lack of clarity regarding the diversity and evolutionary trajectories of modern workload schedulers. We propose a cross-layer taxonomy comprising three categories: OS process scheduling, cluster job scheduling, and big-data scheduling. Through algorithmic feature analysis and historical comparative study, we systematically characterize the design rationales, optimization objectives, and technological evolution of these schedulers, uncovering shared design patterns across local and distributed environments. Our key contribution is the first unified classification framework, which identifies three fundamental differentiating dimensions: resource abstraction granularity, scheduling timing, and feedback mechanism. Based on this analysis, we distill general-purpose scheduling design principles targeting heterogeneity, scalability, and QoS guarantees. The study provides both theoretical foundations and practical guidance for scheduler selection, cross-layer coordination optimization, and next-generation scheduler architecture design.
This work investigates the performance trade-offs of non-adaptive strategies in stochastic load balancing. It proposes a two-stage model: in the first stage, each job reserves up to $k$ machines based on the task size distribution; in the second stage, after observing the actual job size, it is assigned to one of the reserved machines to minimize the expected makespan. The paper establishes, for the first time in this setting, a “power of two choices” theory, showing that under identical machines, reserving just two machines per job suffices to achieve a constant-factor approximation to the omniscient optimal solution. For related machines, it provides an $O(\log m / \log \log m)$-approximation and a bicriteria constant-factor approximation, and further proves that with 2-reservation, one can approximate the adaptively optimal solution.
This study addresses the challenges microservices face in dynamic environments—such as load fluctuations, network variations, and failures—which hinder the coordination of scaling, routing, and repair strategies. The work presents the first taxonomy for adaptive microservice management tailored to dynamic settings, systematically reviewing 84 systems and 13 evaluation artifacts across four dimensions: control placement, dynamic modeling, adaptation strategies, and evaluation evidence. It identifies critical limitations in existing approaches, particularly incomplete modeling of dynamics and insufficient evaluation fidelity, underscoring the importance of high-fidelity evaluation for realizing performance gains. The paper further outlines promising future directions, including cross-layer coordination, telemetry-driven control abstractions, and safe learning-based control, offering a structured roadmap for subsequent research.
This study addresses performance optimization of load balancing strategies in multi-datacenter cloud environments. Using the Cloud Analyst platform, it systematically evaluates Round Robin, Equally Spread, and Throttled algorithms under centralized versus distributed resource architectures and dynamic workloads, measuring response latency and operational cost. Key contributions include: (1) empirical validation that geographic resource distribution significantly impacts latency; (2) in single-datacenter settings, Round Robin achieves marginally lower latency, whereas in cross-datacenter scenarios, Equally Spread and Throttled—particularly when coordinated—yield the lowest average response time (up to 32% reduction) and minimal resource scheduling overhead (27% cost reduction); and (3) demonstration that this synergy effectively balances service quality and economic efficiency. The findings provide evidence-based guidance for designing adaptive, heterogeneous-cloud-aware load balancing policies.
This study addresses scheduler RPC overload in shared Slurm clusters caused by large-scale Nextflow workloads. We propose a reproducible measurement protocol and unified timing attribution method to quantitatively evaluate performance trade-offs across four deployment strategies. By benchmarking Nextflow, HyperQueue, and Flux, we characterize the Pareto frontier between wall-clock time and RPC overhead. Our results demonstrate that Flux significantly reduces scheduling load, establishing it as an optimal solution for high-performance computing sites. These findings provide quantitative guidance for selecting deployment strategies that maintain scheduler stability while preserving user experience in multi-tenant HPC environments.
该研究通过ContinuumBench基准解决了云-边缘环境中服务放置与自动扩展联合评估的问题,采用控制实验方法比较了不同控制器在多种条件下的表现。
This work addresses the technical and behavioral challenges of transitioning from node-exclusive to resource-aware scheduling in production-grade heterogeneous HPC systems, a shift that risks disrupting established scientific workflows. To enable seamless, non-disruptive migration, the authors propose a collaborative operational framework integrating a time-bound compatibility layer, observability-driven feedback mechanisms, and targeted user guidance. Built upon Slurm’s TRES resource model, the approach combines runtime compatibility support, job queue monitoring, and user behavior analysis to preserve workflow continuity while substantially improving scheduling efficiency. Empirical results demonstrate dramatic reductions in median queue wait times—from 277 minutes to under 3 minutes for CPU jobs and from 81 minutes to 3.4 minutes for GPU jobs—alongside high long-term adoption rates among users who embraced the new submission paradigm.
This work addresses the challenges of workflow task composition in high-throughput, petabyte-scale data processing environments, where resource heterogeneity and execution overhead significantly impact performance. The authors propose a hybrid task composition strategy that dynamically balances task independence against execution grouping, formulated within a multi-objective optimization framework to achieve Pareto-optimal trade-offs among throughput, I/O cost, and CPU efficiency. Leveraging workflow DAG modeling and high-dimensional parameter space simulation, the approach enables policy-driven automated synthesis of workflows. Experimental results demonstrate that the proposed strategy achieves up to a 3.8× improvement in throughput and reduces network overhead by as much as 14.9× compared to baseline methods, offering a scalable workflow synthesis framework for extreme-scale scientific computing.
This study investigates the scalability and performance of process and thread schedulers under memory-intensive workloads in multi-core shared-memory systems, focusing on a 3D tensor row-sorting task. The authors design and evaluate several scheduling strategies: on the thread side, an AIMD-based adaptive chunking mechanism inspired by TCP congestion control is introduced, coupled with exponential weighted moving average to dynamically adjust concurrency; on the process side, a bounded prolific/collective model is employed alongside one-to-one, one-to-many, and many-to-many pipelined communication patterns to enable flexible task distribution. Experimental results on a 24-core x86-64 platform demonstrate that thread-level scheduling consistently outperforms process-level scheduling, with dynamic and guided strategies achieving the best performance, while the many-to-many pipeline exhibits superior scalability for large-scale tasks.