PulseInfer: I/O-Centric Sparse KV Cache Offloading for Efficient Long-Context LLM Decoding

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the CPU-GPU I/O bottleneck and transfer fragmentation arising from sparse KV cache offloading during long-context LLM decoding by proposing an I/O-centric architecture. The core methodology introduces interruptible hierarchical scheduling to mask variable latency, an I/O-adaptive admission mechanism for load control, and SoloHead, a sparse selection and aggregation engine that consolidates fragmented transfers. System-level optimizations are implemented within the SGLang framework. Experimental results demonstrate that the proposed approach achieves up to a 4.7× improvement in decoding throughput over baselines and reduces time per output token (TPOT) by 76%, all while maintaining near-lossless accuracy.
📝 Abstract
Long-context LLM serving is increasingly bottlenecked by decode, where large KV caches limit batch size and underutilize GPUs. Sparse KV cache offloading expands effective capacity by storing most historical KV blocks in CPU DRAM and recalling only selected blocks on demand. However, we find that existing offloading systems shift the bottleneck to CPU-GPU recall I/O: recall volume varies widely across layers, decode steps and requests, while headwise sparse selection fragments recalls into many small PCIe transfers. This paper presents PulseInfer, an I/O-centric sparse KV cache offloading system. PulseInfer hides variable recall latency with interruptible layer-wise scheduling, adapts offloading decisions with IO-Adaptive Offloading Admission, and coalesces fragmented transfers using SoloHead sparse selection and a gather-scatter I/O engine. Implemented on SGLang, PulseInfer improves decode throughput by up to 4.7x over SGLang and 2.6x over the best existing offloading baseline, while reducing TPOT by up to 76% and preserving near-lossless accuracy.
Problem

Research questions and friction points this paper is trying to address.

Long-context LLM
KV cache offloading
Sparse attention
I/O bottleneck
Decoding throughput
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse KV Cache Offloading
I/O-Centric Scheduling
Long-Context LLM Decoding
SoloHead Sparse Selection
Gather-Scatter I/O Engine
🔎 Similar Papers
Q
Qiuyang Zhang
Huazhong University of Science and Technology, Wuhan, China
K
Kai Zhou
Huazhong University of Science and Technology, Wuhan, China
Kai Lu
Kai Lu
Postdoc, Huazhong University of Science and Technology
Distributed storage systemskey-value storageAI storage
H
Haocheng Lu
Huazhong University of Science and Technology, Wuhan, China
Jian Zhou
Jian Zhou
Huazhong University of Science and Technology
High Performance ComputingStorage System
Y
Yuanpeng Su
UCloud Technology Co., Ltd., Shanghai, China
K
Kun Bao
UCloud Technology Co., Ltd., Shanghai, China
J
Jiguang Wan
Huazhong University of Science and Technology, Wuhan, China
Fei Wu
Fei Wu
Huazhong University of Science and Technology, Wuhan, China