Is the GPU Half-Empty or Half-Full? Practical Scheduling Techniques for LLMs

📅 2024-10-23
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses inefficient GPU resource scheduling in large language model (LLM) inference serving. We propose a two-tier cooperative scheduling framework: server-level scheduling for load balancing and service-level scheduling optimized for request latency sensitivity. Our approach introduces a lightweight, deployable dynamic priority queue and a preemptive batching mechanism—requiring no modifications to models, hardware, or underlying inference frameworks—and maintains full compatibility with mainstream LLM serving systems. Evaluated under real production workloads, it reduces average tail latency by 22%, improves GPU utilization by 18%, and increases throughput by 15% over state-of-the-art production-grade scheduling policies. The core contribution is a practical, high-performance scheduling paradigm that achieves significant efficiency gains with minimal implementation overhead, delivering a production-ready resource optimization solution for LLM inference serving.

Technology Category

Planning, Routing, and Scheduling: Planning with Language ModelsMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language Models

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Serving systems for Large Language Models (LLMs) improve throughput by processing several requests concurrently. However, multiplexing hardware resources between concurrent requests involves non-trivial scheduling decisions. Practical serving systems typically implement these decisions at two levels: First, a load balancer routes requests to different servers which each hold a replica of the LLM. Then, on each server, an engine-level scheduler decides when to run a request, or when to queue or preempt it. Improved scheduling policies may benefit a wide range of LLM deployments and can often be implemented as"drop-in replacements"to a system's current policy. In this work, we survey scheduling techniques from the literature and from practical serving systems. We find that schedulers from the literature often achieve good performance but introduce significant complexity. In contrast, schedulers in practical deployments often leave easy performance gains on the table but are easy to implement, deploy and configure. This finding motivates us to introduce two new scheduling techniques, which are both easy to implement, and outperform current techniques on production workload traces.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Resource Allocation
GPU Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decision Method
Resource Allocation
Multi-task Processing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
MIT | Databricks
F
Ferdinand Kossmann
MIT CSAIL
B
Bruce Fontaine
Databricks
D
D. Khudia
Databricks
M
Michael J. Cafarella
MIT CSAIL
Samuel Madden
Samuel Madden
MIT
Computer SystemsDatabase SystemsData ManagementMobile ComputingDistributed Systems