🤖 AI Summary
This work addresses the scheduling challenge of simultaneously achieving KV cache reuse, latency SLO compliance, and GPU resource minimization in large-scale LLM serving for agents. We propose an interference-aware request packing mechanism centered on a compact white-box model that accurately predicts interference-induced latency during both prefill and decode phases. Leveraging these predictions, we design an SLO-aware scheduling algorithm that efficiently consolidates requests onto fewer instances while preserving cache reuse, thereby improving per-GPU throughput. Experimental results demonstrate that our approach reduces GPU hours by 16.8%–24.6% compared to state-of-the-art baselines. Furthermore, deployment in a production cluster comprising over one thousand H20 GPUs achieves a 34.7% reduction in resource overhead while strictly satisfying time-per-output-token (TPOT) targets.
📝 Abstract
Request scheduling is a key challenge in large-scale clusters serving agentic large language model (LLM) workloads. An effective scheduler must preserve key-value cache (KVC) reuse across long, shared prefixes, meet token-level latency service-level objectives (SLOs), and minimize GPU resource footprint. Existing schedulers struggle to reconcile these requirements: request consolidation can sacrifice cache locality and increase prefill/decode interference, compromising both SLO attainment and resource efficiency. We present PackServe, a scheduler designed to reduce resource costs while meeting latency SLOs for agentic LLM serving. PackServe uses compact white-box models to predict latency under prefill/decode interference. Guided by these predictions, it packs requests onto fewer serving instances while preserving KVC reuse and SLO constraints, trading available latency headroom for improved per-instance throughput. Evaluation on 64 NVIDIA H20 GPUs shows that PackServe uses up to 16.8% and 24.6% fewer GPU-hours than state-of-the-art schedulers under 30-ms and 50-ms TPOT targets, respectively, while meeting the target TPOT objectives. PackServe has also been deployed in our production cluster comprising over 1000 GPUs, where it reduces the resource footprint by 34.7% compared with the original production scheduler.