Score
Designs, builds, and evaluates algorithms, placement and scheduling policies, and resource-allocation or autoscaling mechanisms that map workloads (jobs, services, or tasks) to computing resources and configuration parameters to satisfy objectives such as throughput, latency, cost, energy use, or fairness under capacity and SLA constraints. Includes profiling and modeling workloads, formulating optimization problems, and implementing controllers, schedulers, or tuning procedures for placement, scaling, load balancing, and configuration selection.
This paper investigates the joint optimization of server count, scheduling policy, and system architecture under a fixed computational budget to minimize average job response time. Using high-resolution traces from Google Cloud production workloads, we develop a multi-stage server cluster model and systematically compare classical policies—including Join-Idle-Queue (JIQ) and Round-Robin (RR)—against state-of-the-art size-aware schedulers. Our findings reveal: (1) an optimal critical server scale that minimizes response time; (2) in high-parallelism or multi-tier architectures, RR and JIQ significantly outperform conventional size-aware policies; and (3) parallelism degree and architectural design exert greater influence on performance than scheduling algorithm sophistication. Collectively, these results establish a new optimization paradigm wherein “architecture–parallelism” dominates over “algorithmic refinement.”
This paper addresses the lack of systematic optimization for CPU and memory resource allocation during the Release phase of cloud-native DevOps. We propose the first pre-deployment offline performance optimization framework for microservices—distinct from mainstream auto-scaling research focused on the Ops phase. Our approach performs fine-grained resource configuration tuning *before* deployment, thereby mitigating auto-scaling failures caused by suboptimal memory provisioning. Methodologically, we integrate Bayesian optimization, statistical experimental design, and a goal-directed factor screening strategy to balance sampling cost and approximation accuracy. Extensive evaluation on the TeaStore benchmark demonstrates that our pre-deployment optimization significantly improves memory suitability and API-level resource utilization. Moreover, it empirically validates the necessity and context-dependent applicability of factor screening under varying optimization objectives.
This paper addresses the lack of clarity regarding the diversity and evolutionary trajectories of modern workload schedulers. We propose a cross-layer taxonomy comprising three categories: OS process scheduling, cluster job scheduling, and big-data scheduling. Through algorithmic feature analysis and historical comparative study, we systematically characterize the design rationales, optimization objectives, and technological evolution of these schedulers, uncovering shared design patterns across local and distributed environments. Our key contribution is the first unified classification framework, which identifies three fundamental differentiating dimensions: resource abstraction granularity, scheduling timing, and feedback mechanism. Based on this analysis, we distill general-purpose scheduling design principles targeting heterogeneity, scalability, and QoS guarantees. The study provides both theoretical foundations and practical guidance for scheduler selection, cross-layer coordination optimization, and next-generation scheduler architecture design.
This work addresses the challenge of resource allocation in geographically distributed and heterogeneous continuum computing infrastructures, where combinatorial explosion and limited generalization hinder effective deployment. To tackle this, the study introduces, for the first time, the pricing structures commonly found in Software-as-a-Service (SaaS) ecosystems into the resource allocation problem, formulating a unified, price-based representation of the configuration space. The authors propose PRIME, a pricing-aware analysis engine that efficiently searches for cost-optimal deployment configurations satisfying both functional and non-functional constraints. Leveraging synthetic infrastructure topologies and workload generation techniques, the project constructs a comprehensive dataset comprising 9,600 diverse scenarios, demonstrating that the proposed approach achieves both scalability and computational efficiency in complex, heterogeneous environments.
Existing scheduling theory struggles to handle multi-resource job scenarios with continuously distributed resource demands, as it relies on the assumption of finitely many job types—a simplification inconsistent with the high heterogeneity observed in real-world workloads. This work proposes the first family of throughput-optimal scheduling policies for continuous multi-resource job models, encompassing both preemptive and non-preemptive variants. The approach employs an adaptive discretization mechanism that dynamically adjusts granularity based on system load and demand distribution. By integrating throughput-optimal control, distribution-aware scheduling, and queueing optimization, the method achieves theoretical optimality while substantially improving computational efficiency. Experiments demonstrate superior performance over state-of-the-art index-based policies under both parametric distributions and real-world Google Borg traces, attaining industry-leading results.
This study addresses the trade-off between energy efficiency and latency for AI/ML workloads in multi-instance GPU (MIG) environments by proposing a dynamic repartitioning scheduling framework tailored to individual MIG instances. The framework integrates scheduler selection with a reinforcement learning–based dynamic repartitioning strategy, marking the first application of reinforcement learning to MIG reconfiguration decisions. It uncovers optimal GPU partitioning patterns under varying temporal and queue-state conditions, enabling predictive and automatic adjustments. Experimental evaluation using real-world diurnal workload traces from data centers demonstrates that the proposed approach improves the combined metric of energy consumption and task latency by 26%, 31%, and 68% compared to twice-daily repartitioning, static partitioning, and no partitioning schemes, respectively.
This work investigates the performance trade-offs of non-adaptive strategies in stochastic load balancing. It proposes a two-stage model: in the first stage, each job reserves up to $k$ machines based on the task size distribution; in the second stage, after observing the actual job size, it is assigned to one of the reserved machines to minimize the expected makespan. The paper establishes, for the first time in this setting, a “power of two choices” theory, showing that under identical machines, reserving just two machines per job suffices to achieve a constant-factor approximation to the omniscient optimal solution. For related machines, it provides an $O(\log m / \log \log m)$-approximation and a bicriteria constant-factor approximation, and further proves that with 2-reservation, one can approximate the adaptively optimal solution.
This work addresses the lack of a unified and comparable benchmark for fairly evaluating rule-based, learning-based, and large language model (LLM)-driven autoscaling strategies in big data batch processing scenarios. To this end, we propose BatchBench, an open-source, workload-aware benchmarking framework. BatchBench introduces a taxonomy encompassing six representative batch workload types, features a parameterized workload generator whose fidelity is validated via two-sample Kolmogorov–Smirnov tests and Earth Mover’s Distance, and defines a five-dimensional evaluation protocol covering cost, SLA compliance, responsiveness, scaling jitter, and interpretability. Notably, it enables, for the first time, side-by-side comparison of all three autoscaling strategy categories while incorporating LLM inference cost accounting. The framework’s design is complete, and its reference implementation will be open-sourced to establish a standardized experimental foundation for autoscaling research.
This work addresses the challenge of efficient container scheduling in dynamic and heterogeneous edge environments, where balancing service-level objectives (SLOs), resource utilization, and data locality is critical. To this end, the authors propose Scale, a novel framework that, for the first time, unifies SLO constraints, end-to-end latency, and data locality into a single optimization model and leverages policy-based deep reinforcement learning to achieve joint optimization. The approach significantly enhances scheduling performance while maintaining system stability. Experimental evaluation on a large-scale real-world dataset from Huawei Cloud demonstrates that Scale achieves scheduling quality within 1.11–1.15× of the optimal solution obtained by integer linear programming, while accelerating decision-making by up to 99%.