Score
Designs and implements scheduling and queueing mechanisms that allocate system resources (CPU time, I/O bandwidth, slots) among users, groups, or jobs according to configured shares or weights; builds algorithms and queue disciplines that track usage, enforce entitlements, and handle dynamic arrivals while analyzing fairness, efficiency, latency, and isolation properties under realistic and adversarial workloads.
This paper addresses the lack of clarity regarding the diversity and evolutionary trajectories of modern workload schedulers. We propose a cross-layer taxonomy comprising three categories: OS process scheduling, cluster job scheduling, and big-data scheduling. Through algorithmic feature analysis and historical comparative study, we systematically characterize the design rationales, optimization objectives, and technological evolution of these schedulers, uncovering shared design patterns across local and distributed environments. Our key contribution is the first unified classification framework, which identifies three fundamental differentiating dimensions: resource abstraction granularity, scheduling timing, and feedback mechanism. Based on this analysis, we distill general-purpose scheduling design principles targeting heterogeneity, scalability, and QoS guarantees. The study provides both theoretical foundations and practical guidance for scheduler selection, cross-layer coordination optimization, and next-generation scheduler architecture design.
This paper investigates the joint optimization of server count, scheduling policy, and system architecture under a fixed computational budget to minimize average job response time. Using high-resolution traces from Google Cloud production workloads, we develop a multi-stage server cluster model and systematically compare classical policies—including Join-Idle-Queue (JIQ) and Round-Robin (RR)—against state-of-the-art size-aware schedulers. Our findings reveal: (1) an optimal critical server scale that minimizes response time; (2) in high-parallelism or multi-tier architectures, RR and JIQ significantly outperform conventional size-aware policies; and (3) parallelism degree and architectural design exert greater influence on performance than scheduling algorithm sophistication. Collectively, these results establish a new optimization paradigm wherein “architecture–parallelism” dominates over “algorithmic refinement.”
This work addresses inefficient GPU resource scheduling in large language model (LLM) inference serving. We propose a two-tier cooperative scheduling framework: server-level scheduling for load balancing and service-level scheduling optimized for request latency sensitivity. Our approach introduces a lightweight, deployable dynamic priority queue and a preemptive batching mechanism—requiring no modifications to models, hardware, or underlying inference frameworks—and maintains full compatibility with mainstream LLM serving systems. Evaluated under real production workloads, it reduces average tail latency by 22%, improves GPU utilization by 18%, and increases throughput by 15% over state-of-the-art production-grade scheduling policies. The core contribution is a practical, high-performance scheduling paradigm that achieves significant efficiency gains with minimal implementation overhead, delivering a production-ready resource optimization solution for LLM inference serving.
To address resource supply-demand imbalances—manifesting as shortages and surpluses—across globally distributed heterogeneous computing clusters, this paper proposes a resource rationing mechanism grounded in real-world market economics. Methodologically, it introduces a periodic simulated-clock auction framework integrating utilization-driven reserve-price setting, long-term resource quota modeling, and supply-demand equilibrium pricing, enabling dynamic price signals to guide users’ autonomous job placement decisions. Its key contribution lies in being the first to systematically embed microeconomic market mechanisms into large-scale distributed resource allocation, replacing static quota or immediate-scheduling paradigms. Evaluated on the Google experimental market, the mechanism significantly incentivizes user migration toward underutilized clusters: resource utilization variance decreases by 32%, and shortage rate drops by 41%. These results empirically validate that price-based incentives can effectively drive system-level behavioral optimization and achieve global resource equilibrium.
This paper addresses online scheduling in a parallel queue system with multiple job classes and multiple servers, where rewards are unknown, dynamically stochastic, and exhibit a bilinear structure. The objective is to jointly maximize cumulative reward and minimize job holding delay (i.e., holding cost), while ensuring system stability—namely, throughput optimality and bounded queue lengths. We propose the first distributed algorithm integrating three key components: (i) dynamic learning of bilinear bandit rewards, (ii) weighted proportional-fair scheduling, and (iii) marginal-cost correction. Theoretically, the algorithm achieves a sublinear regret bound and guarantees bounded expected queue lengths. Empirically, it significantly outperforms existing baselines in both cumulative reward and average delay across computational service and online platform scenarios.
This study investigates the scalability and performance of process and thread schedulers under memory-intensive workloads in multi-core shared-memory systems, focusing on a 3D tensor row-sorting task. The authors design and evaluate several scheduling strategies: on the thread side, an AIMD-based adaptive chunking mechanism inspired by TCP congestion control is introduced, coupled with exponential weighted moving average to dynamically adjust concurrency; on the process side, a bounded prolific/collective model is employed alongside one-to-one, one-to-many, and many-to-many pipelined communication patterns to enable flexible task distribution. Experimental results on a 24-core x86-64 platform demonstrate that thread-level scheduling consistently outperforms process-level scheduling, with dynamic and guided strategies achieving the best performance, while the many-to-many pipeline exhibits superior scalability for large-scale tasks.
Existing scheduling theory struggles to handle multi-resource job scenarios with continuously distributed resource demands, as it relies on the assumption of finitely many job types—a simplification inconsistent with the high heterogeneity observed in real-world workloads. This work proposes the first family of throughput-optimal scheduling policies for continuous multi-resource job models, encompassing both preemptive and non-preemptive variants. The approach employs an adaptive discretization mechanism that dynamically adjusts granularity based on system load and demand distribution. By integrating throughput-optimal control, distribution-aware scheduling, and queueing optimization, the method achieves theoretical optimality while substantially improving computational efficiency. Experiments demonstrate superior performance over state-of-the-art index-based policies under both parametric distributions and real-world Google Borg traces, attaining industry-leading results.
This work addresses the lack of predictable latency and formal guarantees for workloads in multi-tenant GPU clusters, where existing systems rely on heuristic policies without theoretical foundations. The paper presents the first formulation of cluster admission control as a multi-class, multi-resource queueing network, integrating an M/G/k queueing model with a vector bin-packing reduction. It introduces an effective service capacity metric, \(k_{\text{eff}}\), to identify bottleneck dimensions and distinguish schedulable from infeasible loads. Theoretically, it establishes the existence of a schedulable subset, proves that waiting time follows an \(O(1/(1-\rho))\) scaling law, and shows that optimal admission ordering under multidimensional resources is NP-hard. Experiments based on Kueue validate the accuracy of Little’s Law and demonstrate that Erlang-C provides a conservative estimate of waiting times.
This study addresses resource fragmentation in container scheduling and the challenges of operator scheduling dependencies and memory safety in multi-tenant model serving. We propose SliceScheduler, a system that leverages global mapping graph abstraction and incremental simulation to perceive cluster states in real time and enable dynamic operator-level scheduling. This approach achieves practical fine-grained GPU resource multiplexing while maintaining Service Level Agreement (SLA) guarantees for the first time. Experimental results demonstrate that SliceScheduler increases token throughput by 1.10× to 2.29× while keeping SLA violation rates below 9%, effectively validating the superiority of operator-level scheduling in multi-tenant scenarios.
This study addresses a critical limitation in classical queueing analysis—its frequent neglect of preemption overhead—which hinders accurate assessment of stability and response time in preemptive scheduling systems. Focusing on the M/G/1 queue with preemption overhead, this work investigates class-based preemptive priority scheduling and presents the first exact analysis of response time distributions for such systems. By introducing a novel theoretical construct termed “task joint transform,” which integrates Laplace transforms with stochastic process techniques, the authors derive recursive formulas for the Laplace transforms of response times for tasks of arbitrary classes. This framework enables closed-form computation of all response time moments, clearly elucidates the performance impact of preemption overhead, and establishes a general analytical foundation extendable to broader scheduling overhead models.