Score
Design, build, and analyze analytical and simulation queueing models and management mechanisms that predict and control system capacity, throughput, waiting time, and stability; implement and evaluate scheduling, priority, and reserve-based allocation policies and integrate those policies with runtime or resource-management systems.
This work addresses inefficient GPU resource scheduling in large language model (LLM) inference serving. We propose a two-tier cooperative scheduling framework: server-level scheduling for load balancing and service-level scheduling optimized for request latency sensitivity. Our approach introduces a lightweight, deployable dynamic priority queue and a preemptive batching mechanism—requiring no modifications to models, hardware, or underlying inference frameworks—and maintains full compatibility with mainstream LLM serving systems. Evaluated under real production workloads, it reduces average tail latency by 22%, improves GPU utilization by 18%, and increases throughput by 15% over state-of-the-art production-grade scheduling policies. The core contribution is a practical, high-performance scheduling paradigm that achieves significant efficiency gains with minimal implementation overhead, delivering a production-ready resource optimization solution for LLM inference serving.
Existing scheduling theory struggles to handle multi-resource job scenarios with continuously distributed resource demands, as it relies on the assumption of finitely many job types—a simplification inconsistent with the high heterogeneity observed in real-world workloads. This work proposes the first family of throughput-optimal scheduling policies for continuous multi-resource job models, encompassing both preemptive and non-preemptive variants. The approach employs an adaptive discretization mechanism that dynamically adjusts granularity based on system load and demand distribution. By integrating throughput-optimal control, distribution-aware scheduling, and queueing optimization, the method achieves theoretical optimality while substantially improving computational efficiency. Experiments demonstrate superior performance over state-of-the-art index-based policies under both parametric distributions and real-world Google Borg traces, attaining industry-leading results.
This paper addresses online scheduling in a parallel queue system with multiple job classes and multiple servers, where rewards are unknown, dynamically stochastic, and exhibit a bilinear structure. The objective is to jointly maximize cumulative reward and minimize job holding delay (i.e., holding cost), while ensuring system stability—namely, throughput optimality and bounded queue lengths. We propose the first distributed algorithm integrating three key components: (i) dynamic learning of bilinear bandit rewards, (ii) weighted proportional-fair scheduling, and (iii) marginal-cost correction. Theoretically, the algorithm achieves a sublinear regret bound and guarantees bounded expected queue lengths. Empirically, it significantly outperforms existing baselines in both cumulative reward and average delay across computational service and online platform scenarios.
Telecommunications engineering graduate students often lack foundational knowledge in probability theory and stochastic processes, hindering their mastery of teletraffic analysis. Method: This work develops a balanced theoretical–practical pedagogical framework centered on classical queueing models (e.g., M/M/1, M/G/1) and stochastic processes (e.g., Poisson processes, Markov chains, steady-state analysis), integrated with contemporary telecommunications use cases—including traffic modeling, resource allocation, and flow management—and reinforced through numerical simulation exercises. A structured background remediation module addresses prerequisite gaps. Contribution/Results: The resulting textbook has been adopted as a core course resource at multiple universities worldwide, demonstrably enhancing students’ practical competencies in performance modeling and optimization of communication systems.
This work addresses real-time dynamic scheduling in large-scale multi-class call centers, modeled under the Halfin-Whitt heavy-traffic regime to minimize the expected total cost—arising from customer waiting and abandonment—over a finite horizon. We propose the first deep neural network-based simulation optimization framework that approximates high-dimensional (up to 500 classes) queueing systems as diffusion control problems and optimizes policies via end-to-end training. Our method integrates stochastic optimal control theory, diffusion approximations of queueing dynamics, and simulation-based policy gradient estimation. Experiments on real-world data demonstrate that our approach significantly outperforms existing benchmarks while exhibiting strong scalability—overcoming the longstanding dimensionality bottleneck in high-dimensional queueing control.
This study addresses a critical limitation in classical queueing analysis—its frequent neglect of preemption overhead—which hinders accurate assessment of stability and response time in preemptive scheduling systems. Focusing on the M/G/1 queue with preemption overhead, this work investigates class-based preemptive priority scheduling and presents the first exact analysis of response time distributions for such systems. By introducing a novel theoretical construct termed “task joint transform,” which integrates Laplace transforms with stochastic process techniques, the authors derive recursive formulas for the Laplace transforms of response times for tasks of arbitrary classes. This framework enables closed-form computation of all response time moments, clearly elucidates the performance impact of preemption overhead, and establishes a general analytical foundation extendable to broader scheduling overhead models.
This paper addresses the time-sliced task allocation problem in a Markovian state machine setting, aiming to jointly minimize the Age of Job Completion (AoJC) and sampling cost while ensuring stability of $N$ user queues. Recognizing that conventional age metrics fail to capture task completion timeliness, we introduce AoJC as a novel performance metric. Building upon this, we propose two state-aware stable scheduling policies that integrate stochastic arrival modeling, Markov chain-based state machine analysis, and Lyapunov drift-plus-penalty theory. Numerical experiments demonstrate that our policies reduce AoJC by an average of 23.6% compared to baselines, maintain strong system stability under tight sampling budgets, and outperform existing approaches in both age reduction and queue stability guarantees.
This study addresses the complex interplay of dynamically evolving customer classes, abandonment behavior, and dynamic prioritization in finite-capacity, multi-server queueing systems. To tackle this challenge, the authors propose a scalable continuous-time Markov chain (CTMC) modeling framework that integrates quasi-birth–death processes, matrix-analytic methods, and Krylov subspace approximations to efficiently compute both conditional and steady-state waiting time distributions for two customer classes. Notably, this work is the first to incorporate dynamic customer-type evolution and reneging into waiting time analysis for such systems. The model’s validity is demonstrated using real-world data from a tertiary referral hospital in Australia, where it successfully quantifies the disparity in waiting times between complex and routine patients, thereby offering actionable, quantitative insights for healthcare operational decision-making.
This study addresses procedural denials in U.S. SNAP benefit applications caused by call center congestion, which undermines applicants’ due process rights. It introduces, for the first time in social welfare services, a queueing model incorporating redialing and abandonment behaviors, and develops a performance evaluation framework based on fluid approximation and steady-state analysis. This approach corrects the systematic underestimation of system load inherent in traditional Erlang-A models. Calibrated with real-world call data disclosed through court proceedings, the model reveals hidden arbitrariness in service accessibility arising from dynamic interactions between system capacity and demand fluctuations. The proposed framework enables ex ante evaluation of system design, simulation of policy interventions, and provides a quantitative basis for assessing whether applicants have received a meaningful opportunity to access benefits.
This paper investigates stability and heavy-traffic delay optimality for parallel single-server load balancing systems with heterogeneous service rates, under periodic queue-length observations every $T$ time units; the central dispatcher bases decisions solely on the most recent scaled queue-length ordering and server rates. We propose a general class of scheduling policies that jointly leverage scaled ordering and rate awareness. For the first time, we derive necessary and sufficient conditions for system stability under such policies. Furthermore, we establish sufficient conditions for heavy-traffic delay optimality and prove that, in the heavy-traffic limit, the scaled queue-length vector converges weakly to a deterministic vector multiplied by an exponential random scaling factor. Our analysis integrates stochastic process theory, modeling of periodic information updates, and heavy-traffic scaling limit techniques. This work provides the first rigorous stability criterion and delay optimality guarantee for load balancing in heterogeneous systems operating under limited, periodically updated state information.