PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs

📅 2026-06-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the performance bottleneck caused by KV cache movement during long-context large language model (LLM) inference, particularly under mixed sequence lengths and low token activity, where existing scheduling strategies struggle to efficiently utilize consumer-grade GPU resources. The paper proposes PersistentKV, a decoding scheduler for paged KV caches that integrates page-aware scheduling with adaptive mechanisms for the first time. By mapping tasks according to KV head groups, reusing K/V blocks, leveraging native page table support, and maintaining compact work queues, PersistentKV enables fine-grained, non-empty-task-driven scheduling. Experiments on an RTX 3060 demonstrate throughput improvements of 1.06–1.40× over FlashInfer, with no performance degradation in boundary cases, underscoring the critical impact of task assignment on LLM serving efficiency.
📝 Abstract
Autoregressive large language model (LLM) serving is increasingly limited by key-value (KV) cache movement rather than dense matrix multiplication. Modern paged-attention systems reduce KV-cache fragmentation and mature kernels such as FlashInfer provide highly optimized native-paged decode attention. However, the best single-kernel implementation is not always the best serving schedule: low-active long-context decode can under-utilize commodity GPUs, while mixed sequence lengths introduce a tension between many exact-length launches and coarse padded batches. We present PersistentKV, a native block-table decode attention engine and page-aware scheduling study for grouped-query attention (GQA). PersistentKV maps work by KV-head group, is designed to reuse K,V tiles across grouped query heads, supports native page tables, and adds a compact workqueue schedule that executes only non-empty row-KV-head-sequence-split tasks. On an RTX 3060 with FP16, page size 16, Hq=32, Hkv=8, d=128, and identical correctness tolerance against FlashInfer, a calibrated adaptive policy selects FlashInfer for small active batches, PersistentKV sequence splitting for B1 long-context steps, and PersistentKV workqueue scheduling for B8 long-context steps. With thresholds and split counts fixed on calibration traces, one held-out trace seed improves synchronized wall throughput by 1.063-1.265x on B8 bimodal, uniform, and Zipf-like workloads and by 1.399x on a B1 bucketed trace. On the B4 bimodal boundary case, the policy avoids the PersistentKV regression by selecting FlashInfer. These results identify a concrete systems niche for adaptive page-aware decode scheduling and show that work assignment, not only attention math, is a decisive serving-system variable.
Problem

Research questions and friction points this paper is trying to address.

KV cache
long-context LLM serving
decode scheduling
commodity GPUs
paged attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

PersistentKV
page-aware scheduling
KV-cache reuse
grouped-query attention
adaptive decode scheduling