Score
Designs and analyzes resource provisioning and scheduling configurations for service or processing systems to ensure stable operation and required throughput, including deriving maximum sustainable arrival rates and characterizing the stability region under heterogeneous resources. Builds analytical or simulation models to compute throughput capacity for given skill/resource assignments and to quantify capacity gains from collaboration, guiding provisioning, admission control, and task-assignment decisions.
Existing scheduling theory struggles to handle multi-resource job scenarios with continuously distributed resource demands, as it relies on the assumption of finitely many job types—a simplification inconsistent with the high heterogeneity observed in real-world workloads. This work proposes the first family of throughput-optimal scheduling policies for continuous multi-resource job models, encompassing both preemptive and non-preemptive variants. The approach employs an adaptive discretization mechanism that dynamically adjusts granularity based on system load and demand distribution. By integrating throughput-optimal control, distribution-aware scheduling, and queueing optimization, the method achieves theoretical optimality while substantially improving computational efficiency. Experiments demonstrate superior performance over state-of-the-art index-based policies under both parametric distributions and real-world Google Borg traces, attaining industry-leading results.
This work addresses the technical and behavioral challenges of transitioning from node-exclusive to resource-aware scheduling in production-grade heterogeneous HPC systems, a shift that risks disrupting established scientific workflows. To enable seamless, non-disruptive migration, the authors propose a collaborative operational framework integrating a time-bound compatibility layer, observability-driven feedback mechanisms, and targeted user guidance. Built upon Slurm’s TRES resource model, the approach combines runtime compatibility support, job queue monitoring, and user behavior analysis to preserve workflow continuity while substantially improving scheduling efficiency. Empirical results demonstrate dramatic reductions in median queue wait times—from 277 minutes to under 3 minutes for CPU jobs and from 81 minutes to 3.4 minutes for GPU jobs—alongside high long-term adoption rates among users who embraced the new submission paradigm.
Financial institutions face capacity planning and job scheduling challenges in hybrid cloud and on-premise grid environments, where both resource requirements and execution durations exhibit dual uncertainty. Method: This paper proposes a co-optimization framework that jointly minimizes resource provisioning while maximizing service quality—specifically, on-time completion rate. Innovatively, it is the first to jointly model resource and duration uncertainty within capacity planning, employing a constraint programming framework based on paired sampling that integrates deterministic estimation with stochastic sampling for efficient approximate optimization. Contribution/Results: Experiments demonstrate that the method significantly reduces peak resource demand compared to manual scheduling, while maintaining a high on-time completion rate—validating its effectiveness in balancing these conflicting objectives under uncertainty.
To address resource supply-demand imbalances—manifesting as shortages and surpluses—across globally distributed heterogeneous computing clusters, this paper proposes a resource rationing mechanism grounded in real-world market economics. Methodologically, it introduces a periodic simulated-clock auction framework integrating utilization-driven reserve-price setting, long-term resource quota modeling, and supply-demand equilibrium pricing, enabling dynamic price signals to guide users’ autonomous job placement decisions. Its key contribution lies in being the first to systematically embed microeconomic market mechanisms into large-scale distributed resource allocation, replacing static quota or immediate-scheduling paradigms. Evaluated on the Google experimental market, the mechanism significantly incentivizes user migration toward underutilized clusters: resource utilization variance decreases by 32%, and shortage rate drops by 41%. These results empirically validate that price-based incentives can effectively drive system-level behavioral optimization and achieve global resource equilibrium.
This work addresses the challenge of resource allocation in geographically distributed and heterogeneous continuum computing infrastructures, where combinatorial explosion and limited generalization hinder effective deployment. To tackle this, the study introduces, for the first time, the pricing structures commonly found in Software-as-a-Service (SaaS) ecosystems into the resource allocation problem, formulating a unified, price-based representation of the configuration space. The authors propose PRIME, a pricing-aware analysis engine that efficiently searches for cost-optimal deployment configurations satisfying both functional and non-functional constraints. Leveraging synthetic infrastructure topologies and workload generation techniques, the project constructs a comprehensive dataset comprising 9,600 diverse scenarios, demonstrating that the proposed approach achieves both scalability and computational efficiency in complex, heterogeneous environments.
This work addresses the challenge of effectively capturing and exploiting short-term scheduling flexibility while guaranteeing long-term worst-case service for multi-task streams. The authors propose a state-based scheduling framework that models worst-case service guarantees as dynamically updatable states and enforces schedulability by constraining state transitions to remain within a well-defined schedulable polytope. For the first time, they fully characterize this schedulable polytope, employ min-plus algebra for an efficient service model representation, and introduce “hyperbolic service” as a novel, dynamically extensible service form. The proposed approach significantly reduces scheduling decision complexity while strictly preserving quality-of-service guarantees, thereby enhancing practical deployability.
This work addresses the challenge of ensuring service continuity for multi-stage industrial workflows in B5G/6G networks, where conventional per-request QoS mechanisms fall short. To overcome this limitation, the authors propose a capability-aware collaborative planning framework that proactively exposes sustainable QoS capabilities within a finite network planning window. Industrial applications leverage this foresight to map workflow phases and submit demand trajectories, enabling workflow-level, forward-looking joint evaluation and dynamic coordination updates. By integrating network capability modeling, demand mapping, and adaptive coordination, the approach transcends traditional request-granularity constraints. Experimental validation on a real B5G system and large-scale simulations demonstrates that the proposed method significantly enhances service continuity, reduces request rejection rates, and substantially improves workflow completion rates under high network loads.
This work addresses real-time task blocking and latency in low- to mid-Earth orbit distributed satellite constellations caused by stringent onboard resource constraints. To this end, the authors propose the Resource-Aware Task Allocator (RATA), a dynamic scheduling framework built upon a single-layer tree network (SLTN) architecture that jointly incorporates real-time satellite resource states, task arrival rates, and eclipse-induced energy–communication coupling constraints. The study presents the first quantitative characterization of the nonlinear relationship between constellation scale and system performance, identifying CPU availability as the primary bottleneck for task blocking and pinpointing the critical threshold at which system performance transitions from gradual degradation to abrupt collapse. Experimental results demonstrate robust energy supply under solar-aware scheduling; however, CPU bottlenecks intensify rapidly with constellation expansion, thereby establishing a practical upper bound on scalable constellation size.