Score
Design and evaluate placement algorithms and systems that map workloads or tasks onto hosts/resources to optimize communication-related objectives; these solutions incorporate network-aware metrics, topology and multi-bandwidth constraints, and routing information to minimize inter-task communication cost and congestion. They also implement adaptive placement or migration policies that select hosts based on current traffic conditions and routing-informed measurements.
This paper addresses the task scheduling challenge in geo-distributed computing, arising from network heterogeneity, heterogeneous resource pricing, and imbalanced computational capacity. It systematically surveys scheduling techniques across four paradigms: cloud, edge, cloud–edge collaboration, and high-performance computing (HPC). The study introduces the first unified taxonomy covering all four environments, grounded in three core objectives—performance, fairness, and fault tolerance—and identifies cross-cutting challenges including cross-domain latency-sensitive scheduling, multi-regional cost optimization, and elastic fault tolerance. Through bibliometric analysis and qualitative comparative evaluation, it classifies and assesses state-of-the-art approaches—including multi-objective optimization, game-theoretic models, reinforcement learning, and heuristic algorithms. The work traces the technical evolution of scheduling research and proposes six future directions: AI-native schedulers, carbon-aware scheduling, among others—thereby providing theoretical foundations and practical guidance for building adaptive, sustainable distributed scheduling systems.
To address high content delivery costs and latency in dynamic low-earth-orbit (LEO)/medium-earth-orbit (MEO) satellite networks, this paper proposes a trajectory-aware multi-objective replica placement method. It jointly optimizes end-to-end latency, inter-satellite transmission, storage, and replica migration costs by modeling satellite orbital dynamics and spatiotemporal user request distributions, enabling adaptive deployment across geostationary (GEO), LEO, MEO, and hybrid constellations. The key innovations include embedding orbital mechanics into the optimization framework, designing a lightweight trajectory-driven algorithm, and validating the approach via real-world traffic traces and a prototype system. Experiments demonstrate that, compared to baseline methods, the proposed solution reduces transmission latency by 32.7%, decreases bandwidth overhead by 28.4%, and constrains replica migration cost within acceptable bounds—thereby significantly improving content delivery efficiency and resource utilization, especially in remote regions.
This work addresses the challenge of resource allocation in geographically distributed and heterogeneous continuum computing infrastructures, where combinatorial explosion and limited generalization hinder effective deployment. To tackle this, the study introduces, for the first time, the pricing structures commonly found in Software-as-a-Service (SaaS) ecosystems into the resource allocation problem, formulating a unified, price-based representation of the configuration space. The authors propose PRIME, a pricing-aware analysis engine that efficiently searches for cost-optimal deployment configurations satisfying both functional and non-functional constraints. Leveraging synthetic infrastructure topologies and workload generation techniques, the project constructs a comprehensive dataset comprising 9,600 diverse scenarios, demonstrating that the proposed approach achieves both scalability and computational efficiency in complex, heterogeneous environments.
To address low task processing efficiency, inter-satellite link congestion, and imbalanced resource loading in Low Earth Orbit (LEO) satellite networks, this paper proposes a dynamic multi-region collaborative scheduling framework. The framework integrates three key components: (1) a dynamic adaptive multi-region partitioning mechanism; (2) genetic algorithm (GA)-based regional topology reconfiguration; and (3) a multi-agent deep deterministic policy gradient (MA-DDPG)-driven task splitting and cross-domain offloading strategy. By jointly optimizing region partitioning, routing selection, and computational offloading, the framework enables tight coupling and coordinated management of communication and computing resources. Experimental results demonstrate that the proposed approach significantly outperforms baseline methods in terms of task latency, per-task energy consumption, and task completion rate. System resource utilization improves by 23.6%, validating the framework’s effectiveness and scalability in large-scale, highly dynamic LEO networks.
To address slow response to dynamic traffic demands and excessive topology reconfiguration oscillations in reconfigurable data centers, this paper proposes a batched dynamic graph scheduling algorithm based on incremental matching. It is the first work to introduce dynamic graph algorithms into optical circuit-switched network scheduling, modeling edge-disjoint matchings while explicitly capturing spatiotemporal locality of traffic to avoid full recomputation. We design six efficient batched algorithms and evaluate them on 176 synthetic and 39 real-world traffic traces. Compared to static approaches, our method reduces runtime by 92%, decreases configuration changes by 87%, and incurs zero loss in matching weight—achieving millisecond-scale responsiveness and high configuration stability. The core contribution lies in the principled integration of dynamic graph theory with optical switching characteristics, jointly optimizing update efficiency, reconfiguration stability, and matching quality.
Existing load-balancing mechanisms struggle to effectively exploit path diversity in low-diameter topologies such as Dragonfly and Slim Fly, often relying on proprietary hardware or lacking adaptivity. This work proposes Spritz, a sender-based, general-purpose Ethernet load-balancing framework that, for the first time, enables topology-aware adaptive routing without requiring additional hardware support. Spritz integrates two complementary algorithms—Spritz-Scout and Spritz-Spray—that leverage ECN feedback, packet truncation, and timeout signals for efficient path probing and selection, augmented by a caching mechanism to enhance performance. Large-scale simulations at the thousand-node level demonstrate that Spritz reduces flow completion times by up to 1.8× compared to ECMP and UGAL-L under normal conditions, and achieves up to a 25.4× improvement in the presence of link failures.
This work addresses the challenge of real-time request routing in large language model serving, where incoming requests must be scheduled to decoding nodes under constraints on batch size and KV cache capacity. Existing heuristic approaches struggle to explicitly balance the trade-off between latency and throughput. To overcome this limitation, the paper introduces a novel framework that integrates multi-objective optimization with online linear programming. Request admission decisions are made by comparing SLO-weighted rewards against dual shadow prices, while a warm-started first-order projected gradient method efficiently tracks dynamic dual variables, enabling millisecond-scale, interpretable, and tunable scheduling. Experiments on the Vidur simulation platform demonstrate that the proposed approach consistently outperforms baseline methods across diverse SLO configurations, achieving comprehensive improvements in end-to-end latency, time-to-first-token, throughput, and tail performance.
This study addresses the scalability bottleneck in unicast and multicast routing caused by the dual role of IP addresses as both identifiers and locators. It systematically traces the evolution of Internet routing scalability solutions, first articulating the map-and-encap architecture as a unifying paradigm and identifying the essential conditions for its successful deployment. Through historical protocol analysis, architectural comparisons, and conceptual abstraction—encompassing approaches such as BIER and tunnel encapsulation—the work reveals that BGP’s lack of intra-domain egress router topology abstraction is a fundamental limitation. The paper proposes core principles to guide future scalable routing designs, emphasizing the critical roles of locally driven incentives and effective topology abstraction in protocol evolution.
This study addresses energy efficiency optimization in communication networks during low-traffic periods by jointly optimizing network topology design and shortest-path routing. The approach ensures that all traffic demands can be satisfied within the activated subnetwork through dynamically adapted shortest paths. The authors propose, for the first time, a capacitated integer linear programming model that precisely captures dynamic shortest-path routing, complemented by provably effective strengthening constraints to accelerate solution convergence. A tailored column generation algorithm is developed to efficiently handle large-scale instances. Experimental results demonstrate that a simplified strategy—fixing routes and deactivating redundant links—achieves near-optimal performance, while the traffic-oblivious method TOCA exhibits superior efficacy in multi-demand scenarios.