PackServe: SLO-Aware Request Scheduling for Agentic LLM Serving at Scale

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the scheduling challenge of simultaneously achieving KV cache reuse, latency SLO compliance, and GPU resource minimization in large-scale LLM serving for agents. We propose an interference-aware request packing mechanism centered on a compact white-box model that accurately predicts interference-induced latency during both prefill and decode phases. Leveraging these predictions, we design an SLO-aware scheduling algorithm that efficiently consolidates requests onto fewer instances while preserving cache reuse, thereby improving per-GPU throughput. Experimental results demonstrate that our approach reduces GPU hours by 16.8%–24.6% compared to state-of-the-art baselines. Furthermore, deployment in a production cluster comprising over one thousand H20 GPUs achieves a 34.7% reduction in resource overhead while strictly satisfying time-per-output-token (TPOT) targets.
📝 Abstract
Request scheduling is a key challenge in large-scale clusters serving agentic large language model (LLM) workloads. An effective scheduler must preserve key-value cache (KVC) reuse across long, shared prefixes, meet token-level latency service-level objectives (SLOs), and minimize GPU resource footprint. Existing schedulers struggle to reconcile these requirements: request consolidation can sacrifice cache locality and increase prefill/decode interference, compromising both SLO attainment and resource efficiency. We present PackServe, a scheduler designed to reduce resource costs while meeting latency SLOs for agentic LLM serving. PackServe uses compact white-box models to predict latency under prefill/decode interference. Guided by these predictions, it packs requests onto fewer serving instances while preserving KVC reuse and SLO constraints, trading available latency headroom for improved per-instance throughput. Evaluation on 64 NVIDIA H20 GPUs shows that PackServe uses up to 16.8% and 24.6% fewer GPU-hours than state-of-the-art schedulers under 30-ms and 50-ms TPOT targets, respectively, while meeting the target TPOT objectives. PackServe has also been deployed in our production cluster comprising over 1000 GPUs, where it reduces the resource footprint by 34.7% compared with the original production scheduler.
Problem

Research questions and friction points this paper is trying to address.

Agentic LLM Serving
Request Scheduling
SLO-aware
KV Cache Reuse
Resource Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic LLM Serving
Request Scheduling
SLO-Aware
KV Cache Reuse
White-Box Latency Prediction
Z
Zhiyuan Tan
The Chinese University of Hong Kong, Shenzhen
D
Dejiang Zhu
Ant Group
J
Jingzhe Jiang
The Chinese University of Hong Kong, Shenzhen
Y
Yihao Zheng
The Chinese University of Hong Kong, Shenzhen
Y
Yang Tian
Ant Group
T
Tao Wang
Ant Group
Minchen Yu
Minchen Yu
The Chinese University of Hong Kong, Shenzhen
cloud computingserverless computingbig data systemsmachine learning systems