Score
Designs, builds, or analyzes algorithms, architectures, and operational policies that distribute requests, tasks, or data across servers, nodes, or storage to optimize throughput, latency, resource utilization, fairness, or stability. This work covers global and distributed load‑balancing architectures, placement and migration algorithms, dynamic and high‑concurrency strategies, cross‑workload and cross‑tenant integration, and data‑level balancing techniques.
This study addresses performance optimization of load balancing strategies in multi-datacenter cloud environments. Using the Cloud Analyst platform, it systematically evaluates Round Robin, Equally Spread, and Throttled algorithms under centralized versus distributed resource architectures and dynamic workloads, measuring response latency and operational cost. Key contributions include: (1) empirical validation that geographic resource distribution significantly impacts latency; (2) in single-datacenter settings, Round Robin achieves marginally lower latency, whereas in cross-datacenter scenarios, Equally Spread and Throttled—particularly when coordinated—yield the lowest average response time (up to 32% reduction) and minimal resource scheduling overhead (27% cost reduction); and (3) demonstration that this synergy effectively balances service quality and economic efficiency. The findings provide evidence-based guidance for designing adaptive, heterogeneous-cloud-aware load balancing policies.
This work addresses the practical challenges faced by d-choice load balancing in large-scale service systems, where bursty traffic, multi-priority tasks, and information noise significantly degrade load distribution efficiency and system stability. Bridging the gap between theoretical models and real-world deployment, this study systematically extends d-choice balancing along three critical dimensions: burst recovery, support for multiple task priorities, and tolerance to noisy state information. Leveraging large-scale simulations and an analytical framework based on generative models, the authors characterize and validate the policy’s behavior in dynamic, heterogeneous environments. The results demonstrate that the proposed strategy rapidly recovers from traffic bursts, effectively manages tasks of varying priorities, and remains robust under imperfect information, thereby offering a highly resilient scheduling solution for cloud-scale systems.
To address node overload, high operational costs, and poor system stability caused by dynamic heterogeneous resource scheduling in cloud computing, this paper proposes an intelligent load-balancing framework. The method constructs a high-fidelity simulation environment and an abstracted multi-resource model, introducing for the first time a joint resource utilization metric that incorporates VM migration overhead. It establishes a novel three-category taxonomy for schedulers, derives an empirically grounded formula for estimating VM migration traffic, and comparatively evaluates two emerging paradigms: centralized metaheuristic and distributed multi-agent scheduling. Built upon real-world Google cluster traces, the framework integrates live VM migration and realistic workload simulation. Experimental validation on the University of Westminster’s HPC cluster demonstrates a 23.6% improvement in resource utilization, a 31.4% reduction in task latency, and a 27.9% decrease in network migration overhead—significantly enhancing system stability and cost-efficiency.
In distributed systems, joint scheduling of tasks and data across nodes is challenging, and data hotspots cause severe load imbalance. Method: This paper proposes a task-data co-orchestration abstraction supporting bidirectional task and data migration, coupled with a lightweight distributed push-pull mechanism to achieve low communication overhead and high scalability under highly skewed workloads. The approach integrates distributed task scheduling, dynamic data migration, push-pull–based load balancing, and three execution-flow optimization techniques. Contributions/Results: Experiments show up to 2.7× end-to-end performance improvement over state-of-the-art schedulers. Built upon this framework, the TDO-GP system achieves 4.1× average speedup for general-purpose graph processing, significantly enhancing load-balancing efficiency and system throughput in large-scale graph analytics and key-value store workloads.
Traditional load balancing algorithms struggle to cope with dynamic traffic in cloud environments, resulting in high response latency, frequent congestion, and inflexible scheduling. This paper proposes the first reinforcement learning–based closed-loop load balancing framework, integrating Deep Q-Networks (DQN) into a real-time decision-making loop. It jointly models multi-dimensional metrics—including CPU utilization, response time, and queue length—to construct a dynamic state-action space, enabling online policy updates and proactive intervention under non-equilibrium conditions. Evaluated on a microservice cluster simulation, the approach reduces average response time by 37.2%, improves system availability to 99.99%, and decreases overloaded node incidence by 82%. Its core contribution lies in unifying environment awareness, adaptive scheduling, and predictive intervention—thereby significantly enhancing service quality and system resilience under dynamic workloads.
This paper addresses the lack of clarity regarding the diversity and evolutionary trajectories of modern workload schedulers. We propose a cross-layer taxonomy comprising three categories: OS process scheduling, cluster job scheduling, and big-data scheduling. Through algorithmic feature analysis and historical comparative study, we systematically characterize the design rationales, optimization objectives, and technological evolution of these schedulers, uncovering shared design patterns across local and distributed environments. Our key contribution is the first unified classification framework, which identifies three fundamental differentiating dimensions: resource abstraction granularity, scheduling timing, and feedback mechanism. Based on this analysis, we distill general-purpose scheduling design principles targeting heterogeneity, scalability, and QoS guarantees. The study provides both theoretical foundations and practical guidance for scheduler selection, cross-layer coordination optimization, and next-generation scheduler architecture design.
This work addresses the significant synchronization barrier delays and computational inefficiencies in large model inference caused by load imbalance under data parallelism (DP). To mitigate KV cache migration overhead and persistent skew induced by dynamic request patterns, the authors propose BalanceRoute—a family of online routing algorithms that dynamically assign requests to DP workers under millisecond-level scheduling constraints. Key innovations include BR-0, a prediction-free baseline; BR-H, which employs a short planning horizon; a piecewise-linear F-score with discounting to model load safety margins; and an integrated framework featuring two-stage scheduling, a lightweight termination classifier, and KV cache-aware modeling. Evaluated on a 144-NPU cluster, BalanceRoute substantially reduces DP load imbalance compared to vLLM and achieves higher end-to-end throughput on both Azure-2024 and production workloads.
To address service load balancing under multiple resource constraints in cloud environments, this paper proposes an enhanced genetic algorithm integrating high-quality solutions from diverse metaheuristics (e.g., PSO, SA) as the initial population—thereby accelerating convergence and improving solution quality. The method incorporates abstracted resource modeling, fine-grained multi-dimensional load evaluation, and an explicit service migration overhead quantification model to enable cost-aware dynamic scheduling. Experiments on heterogeneous cloud platforms demonstrate that the proposed algorithm reduces average node load by 23.6%, decreases service migration count by 31.4%, and lowers total operational cost by 18.9%, while maintaining system stability and SLA compliance. The core contributions are: (1) a multi-objective optimization framework with explicit migration cost modeling, and (2) empirical validation that multi-strategy initialization significantly enhances the effectiveness of genetic algorithms for cloud workload scheduling.
Scheduling data-intensive workloads in large-scale distributed systems faces challenges including complexity, heterogeneous parallelism, data locality constraints, and multi-dimensional QoS optimization (e.g., timeliness, fault tolerance, energy efficiency). Method: This paper proposes a novel workload classification scheme grounded in data characteristics and service requirements; systematically surveys and structures mainstream scheduling strategies, exposing critical limitations in dynamic adaptability, fine-grained fault tolerance, and energy–QoS co-optimization; and introduces a unified scheduling framework integrating data-locality awareness, elastic parallel scheduling, QoS-tiered guarantees, and energy-aware resource allocation. Contribution/Results: The study establishes a scalable classification paradigm, delivers a clear technology evolution roadmap, and identifies a prioritized list of open research challenges—thereby advancing foundational understanding and guiding future design of intelligent, holistic schedulers for modern distributed data systems.
This paper investigates stability and heavy-traffic delay optimality for parallel single-server load balancing systems with heterogeneous service rates, under periodic queue-length observations every $T$ time units; the central dispatcher bases decisions solely on the most recent scaled queue-length ordering and server rates. We propose a general class of scheduling policies that jointly leverage scaled ordering and rate awareness. For the first time, we derive necessary and sufficient conditions for system stability under such policies. Furthermore, we establish sufficient conditions for heavy-traffic delay optimality and prove that, in the heavy-traffic limit, the scaled queue-length vector converges weakly to a deterministic vector multiplied by an exponential random scaling factor. Our analysis integrates stochastic process theory, modeling of periodic information updates, and heavy-traffic scaling limit techniques. This work provides the first rigorous stability criterion and delay optimality guarantee for load balancing in heterogeneous systems operating under limited, periodically updated state information.