🤖 AI Summary
This study addresses the CPU-GPU I/O bottleneck and transfer fragmentation arising from sparse KV cache offloading during long-context LLM decoding by proposing an I/O-centric architecture. The core methodology introduces interruptible hierarchical scheduling to mask variable latency, an I/O-adaptive admission mechanism for load control, and SoloHead, a sparse selection and aggregation engine that consolidates fragmented transfers. System-level optimizations are implemented within the SGLang framework. Experimental results demonstrate that the proposed approach achieves up to a 4.7× improvement in decoding throughput over baselines and reduces time per output token (TPOT) by 76%, all while maintaining near-lossless accuracy.
📝 Abstract
Long-context LLM serving is increasingly bottlenecked by decode, where large KV caches limit batch size and underutilize GPUs. Sparse KV cache offloading expands effective capacity by storing most historical KV blocks in CPU DRAM and recalling only selected blocks on demand. However, we find that existing offloading systems shift the bottleneck to CPU-GPU recall I/O: recall volume varies widely across layers, decode steps and requests, while headwise sparse selection fragments recalls into many small PCIe transfers.
This paper presents PulseInfer, an I/O-centric sparse KV cache offloading system. PulseInfer hides variable recall latency with interruptible layer-wise scheduling, adapts offloading decisions with IO-Adaptive Offloading Admission, and coalesces fragmented transfers using SoloHead sparse selection and a gather-scatter I/O engine. Implemented on SGLang, PulseInfer improves decode throughput by up to 4.7x over SGLang and 2.6x over the best existing offloading baseline, while reducing TPOT by up to 76% and preserving near-lossless accuracy.