Score
Design and implement decoder runtime components that execute autoregressive decode workflows while managing key-value (KV) cache and attention with page-aware memory management and scheduling. Build adaptive kernel selection, block-table decode-attention implementations, persistent KV tiling, and workqueue scheduling that maps work by KV-head groups to minimize KV-cache movement and to support native page-table layouts.
This work addresses the substantial memory footprint of key-value (KV) cache in large language model (LLM) inference, which significantly increases serving costs and reduces efficiency. For the first time from a systems perspective, it systematically decomposes KV cache optimization into three orthogonal dimensions: execution scheduling (temporal), placement and migration (spatial), and representation preservation (structural), thereby establishing a unified analytical framework termed sKis. This framework integrates core techniques including scheduling policies, memory management, compressed representations, and dynamic migration, revealing synergistic cross-dimensional design mechanisms and their relationships with optimization objectives. The proposed framework provides both theoretical foundations and practical guidance for building efficient LLM serving infrastructure.
Large language models (LLMs) suffer from high memory overhead and inefficiency in long-context and real-time inference scenarios due to KV cache accumulation. To address this, we propose the first holistic, three-tiered KV cache management framework—operating at the token, model, and system levels—that unifies cache selection, quantization, low-rank decomposition, attention sparsification/windowing, and hardware-aware scheduling. We establish a standardized benchmark covering both text and multimodal tasks, release the first open-source repository for KV cache management research (Awesome-KV-Cache-Management), and provide a comprehensive technical taxonomy with empirical comparisons across methods. This work advances the systematization and standardization of KV cache management methodologies, significantly improving inference efficiency and deployment feasibility of LLMs under resource constraints.
This work addresses the challenges of efficiency, scalability, and consistency in distributed management of KV caches for large language model (LLM) services. It proposes the first four-dimensional taxonomy—spanning locality, lifetime, ownership, and storage substrate—to systematically analyze over 30 existing studies, identifying five architectural paradigms: local paging, decoupled pipelining, shared storage, memory pooling, and hybrid hierarchical designs. The study reveals that “ownership” is a key differentiator in distributed KV cache architectures and highlights the absence of seven KV-specific metrics in current evaluation methodologies. Furthermore, it connects these gaps to six critical open problems, including fault tolerance, isolation, hierarchical eviction, and speculative decoding.
This work addresses the performance bottleneck caused by KV cache movement during long-context large language model (LLM) inference, particularly under mixed sequence lengths and low token activity, where existing scheduling strategies struggle to efficiently utilize consumer-grade GPU resources. The paper proposes PersistentKV, a decoding scheduler for paged KV caches that integrates page-aware scheduling with adaptive mechanisms for the first time. By mapping tasks according to KV head groups, reusing K/V blocks, leveraging native page table support, and maintaining compact work queues, PersistentKV enables fine-grained, non-empty-task-driven scheduling. Experiments on an RTX 3060 demonstrate throughput improvements of 1.06–1.40× over FlashInfer, with no performance degradation in boundary cases, underscoring the critical impact of task assignment on LLM serving efficiency.
This work addresses the high memory consumption of key-value (KV) caching in visual autoregressive (VAR) models for image generation, which often exceeds several gigabytes and limits practical deployment. The authors propose a fine-grained, attention-head-specific KV cache compression method that leverages offline calibration to assess each head’s reliance on historical tokens. By integrating attention-score-based head ranking, static pruning, and dynamic cache allocation, the approach enables differentiated compression under a fixed memory budget. Evaluated on the Infinity-2B model, this technique achieves a 2× higher KV cache compression ratio compared to existing methods while preserving or even improving image fidelity, prompt alignment, and human-perceived quality, thereby establishing a new state of the art in VAR model cache compression.
Existing KV cache compression strategies for long-context large language models are rigid and neglect task-specific and layer-wise heterogeneity. Method: This paper proposes a dynamic hierarchical adaptive compression mechanism featuring (1) task-aware dynamic budget allocation, enabling per-layer, online adjustment of retained token counts; and (2) lightweight cache reconfiguration guided by inter- and intra-layer activation pattern analysis, under global and layer-specific budget constraints. The method requires no fine-tuning and performs periodic optimization during inference. Contribution/Results: With only 1.7% of the original KV cache retained, our approach achieves 85% of full-cache LongBench performance; under extreme compression (0.9% cache), it surpasses state-of-the-art methods by 11% accuracy on the Needle-in-a-Haystack task. The method significantly improves the trade-off between inference efficiency and accuracy for long-context processing.
Existing lossless KV cache management approaches overlook the computational efficiency of GPU attention kernels, resulting in high inference latency. This work proposes AsymCache, a system that, for the first time, incorporates GPU computation latency into cache eviction decisions. By integrating multi-segment attention mechanisms, a position-aware recomputation cost model, and adaptive chunked scheduling, AsymCache jointly optimizes cache hit rates and computational efficiency. Compared to state-of-the-art baselines, AsymCache reduces first-token latency by 1.90–2.03× and per-token generation time by 1.62–1.71×, while achieving an average 18.1% reduction in job latency when deployed in the Continuum system.
To address memory and bandwidth bottlenecks induced by KV caching in long-context generation, this paper proposes the first dual-stage separation compression paradigm—distinctly optimizing prefilling and decoding. During prefilling, excessive compression is avoided to preserve contextual understanding; during decoding, a sliding-window re-hitter identification mechanism coupled with adaptive discontinuous memory transfer dynamically retains critical key-value pairs. The method is plug-and-play compatible with mainstream KV compression techniques. Evaluated on LongGenBench, it achieves significant reductions in memory footprint and memory bandwidth consumption compared to prefilling-only compression baselines, while maintaining lossless generation quality and strong generalization across diverse long-context tasks.
This study reveals a pronounced non-monotonic latency behavior in Apple’s Metal Performance Shaders (MPS) backend during autoregressive decoding, challenging the widely held assumption that KV caching universally improves inference efficiency. Through controlled experiments on models including GPT-2, BLOOM, and OPT, the authors systematically compare MPS against CPU and NVIDIA T4 platforms, uncovering latency spikes of up to 21× within specific decoding length ranges on MPS—phenomena unexplained by memory pressure or prefill costs. The work provides the first characterization of the discontinuous latency patterns in MPS decoding, attributing them to backend execution scheduling mechanisms. These findings underscore the critical importance of hardware-aware evaluation for inference optimization and caution that aggregate benchmarking metrics may obscure such pivotal performance anomalies.
This work addresses memory inefficiency and tail latency in static-graph large language model serving, caused by heterogeneous request lengths, asynchronous completion, and fragmented KV cache layouts. The authors propose KV-RM, a runtime system that decouples logical history from physical storage to standardize KV cache movement under fixed decoding interfaces. By employing block-level paging to track active states and designing a coalesced transfer path that aggregates non-contiguous KV mappings into large contiguous blocks, KV-RM aligns with fixed-shape attention kernels. This approach requires neither distant-history summarization nor dynamic scheduling, significantly improving throughput and tail latency under mixed-length request loads, reducing reserved memory overhead, and eliminating bursty latency spikes during production replay.
This work addresses the substantial memory and bandwidth bottlenecks in autoregressive decoding caused by the linear growth of key-value (KV) cache with context length. Existing KV cache eviction methods rely on static heuristics or proxy scores that inadequately estimate each cache entry’s contribution to future token generation, often resulting in significant performance degradation. To overcome this limitation, the authors propose a supervised learning framework that directly optimizes KV cache eviction under a fixed budget by leveraging future attention targets as supervision signals. They further introduce a delayed memory scorer that implicitly guides online pruning using near-future context, eliminating the need for explicit computation of dense attention maps. Evaluated on Qwen3-4B and Qwen3-8B, the method retains 97%–98% of original model performance at aggressive compression rates of 75%–88%, substantially outperforming current baselines.
This work addresses the substantial memory overhead of Key-Value (KV) caching in long-context inference with large language models. It proposes the first low-rank compression framework that treats Keys and Values differently: for the Key cache, it dynamically groups attention heads based on Centered Kernel Alignment (CKA) and adaptively allocates rank budgets; for the Value cache, it employs offline calibration to optimize low-rank decomposition. Evaluated on three instruction-tuned large language models, the method achieves significant compression of Key cache parameters while maintaining competitive accuracy, demonstrating particularly strong performance in multi-head attention architectures.