Score
Designs and implements methods, data structures, and evaluation procedures to compress and manage model context and key-value (KV) caches for transformer-style architectures, including compact vector encoders, symbolic or telegraphic rewrites, entity–relation re-expressions, incremental/revisable dialogue memories, and head-aware/per-thread KV selection and pruning. Builds and analyzes dynamic or elastic context-management policies (e.g., per-head budgets, adaptive pruning, per-thread compression) that preserve salient information for retrieval and generation while minimizing storage and attention overhead.
Large language models (LLMs) suffer from quadratic attention complexity and explosive KV cache memory consumption and latency when processing long contexts (≥32K tokens). To address this, we propose the first comprehensive methodology classification framework for KV cache compression, unifying key techniques—including quantization, pruning, low-rank approximation, sequence grouping, and locality-aware reuse—across both theoretical principles and practical implementation dimensions. Through standardized cross-method evaluation, we characterize the fundamental trade-offs among accuracy, inference speed, and memory footprint. Experimental results demonstrate that, with perplexity degradation under 1%, the optimal compression strategy achieves 2–5× KV cache memory reduction and 1.3–2.1× inference speedup, substantially enhancing the feasibility of deploying LLMs in long-context scenarios.
This work addresses the linear growth of memory and access overhead in KV cache with increasing context length in long-context language models, a challenge exacerbated by the limited compressibility of representations learned during pretraining. The study formalizes KV compressibility as an intrinsic model representation property and introduces KV Compression-Aware Training (KV-CAT), a framework that incorporates a sparsity-inducing masking strategy during continued pretraining to encourage the learning of inherently more compressible representations. By shaping model representations at the source, KV-CAT enhances compatibility with downstream post-hoc compression techniques, significantly improving the trade-off between compression quality and computational budget across retrieval, long-context question answering, and compressed prefix continuation tasks.
KV cache grows linearly with context length, becoming a critical bottleneck in GPU memory capacity and bandwidth during large language model inference. This work presents the first systematic taxonomy of existing KV cache optimization techniques, categorizing them into five classes: cache eviction, compression, hybrid memory management, novel attention mechanisms, and hybrid strategies. The study evaluates these approaches across seven representative deployment scenarios, revealing that no single method universally dominates; instead, optimal choices require careful trade-offs among memory usage, throughput, and accuracy based on context length, hardware constraints, and workload characteristics. The paper further advocates adaptive, multi-stage optimization as a promising direction for future research, offering both theoretical insights and practical guidance for real-world deployment.
To address the quadratic memory overhead of KV caches in large language models (LLMs) with increasing context length, this paper proposes HeadKV-R2, a head-granularity dynamic compression method. HeadKV-R2 introduces the first attention-head-level KV cache compression paradigm, leveraging a context-aware importance scoring mechanism that jointly models retrieval and reasoning capabilities. Guided by question-answering tasks, it quantifies the contribution of each attention head and performs dynamic sub-sampling of KV tokens per head. Evaluated on LongBench and LooGLE benchmarks across Llama-3-8B and Mistral-7B, HeadKV-R2 significantly outperforms layer-wise compression baselines: it retains 97% of full-cache performance while preserving only 1.5% of the original KV cache, and achieves up to 12.7% absolute improvement when KV cache sizes are set to 64 or 128. The implementation is publicly available.
This work addresses the excessive memory footprint of KV caches in Transformer inference, which severely limits long-context processing efficiency. We propose KVzap, an input-adaptive KV cache pruning method that enables fast and high-fidelity dynamic compression during both prefill and decoding stages. Building upon the efficient approximation algorithm of KVzip, KVzap achieves low-overhead, high-accuracy adaptive compression within mainstream inference engines for the first time, effectively overcoming the traditional trade-off between speed and accuracy. Evaluated on Qwen3-8B, Llama-3.1-8B-Instruct, and Qwen3-32B, our method achieves 2–4× KV cache compression with negligible accuracy loss, establishing a new state of the art on the KVpress leaderboard.
To address the exponential growth of Key-Value (KV) cache memory consumption with context length in large language model (LLM) inference—causing severe memory bottlenecks—this work systematically surveys three mainstream optimization paradigms: selective caching, low-bit quantization, and attention compression, exposing their limitations in compute-storage trade-offs, task generalizability, and hardware compatibility. We propose a hybrid optimization framework, a dynamic adaptive caching strategy, and a hardware-software co-design paradigm to overcome the interoperability constraints inherent in single-method approaches. Experiments demonstrate that our method achieves an average 62% reduction in KV cache memory footprint while sustaining less than 1% accuracy degradation, and improves end-to-end inference throughput by 3.1×. The approach provides a scalable, theoretically grounded pathway for efficient deployment of long-context LLMs.
To address the prohibitively large KV cache memory overhead (up to several gigabytes) in long-context inference for Transformer-based large language models (LLMs), this paper proposes a dynamic, scalable, lightweight lossy compression framework. The framework deeply integrates LLM-aware cache characteristics and innovatively combines block-wise partitioning, quantization, and sparsification—co-optimized with attention computation kernels to significantly reduce data movement costs. Compared to state-of-the-art methods, it achieves an average 47% reduction—and up to 83% peak reduction—in KV cache memory usage, with negligible accuracy degradation. Decompression is highly efficient: in certain scenarios, it even accelerates matrix-vector operations, outperforming the native cuBLAS attention kernel in end-to-end latency.
Existing KV cache compression methods often sacrifice semantic recall due to their reliance on attention scores and sliding window mechanisms, and they face challenges in fair performance evaluation. This work proposes LASER-KV, a novel framework that introduces a block-level cumulative budget mechanism and a dynamic protection factor to decouple compression from sliding window interference. By integrating Exact-LSH for semantic-aware compression, LASER-KV achieves high recall without resorting to fixed-size greedy strategies, instead employing hierarchical selection for more rational cache budget allocation. Evaluated on the Babilong benchmark with a 128k context length, LASER-KV improves accuracy by up to 10% over current methods and effectively mitigates performance degradation of 15–30%.
Existing KV cache compression methods employ uniform strategies and homogeneous budgets, failing to accommodate the heterogeneous requirements of different Transformer layers during prefilling and decoding stages. This work proposes PolyKV, the first framework enabling layer-wise heterogeneous KV cache optimization. PolyKV dynamically selects compression strategies per layer based on layer-specific signal analysis and integrates multi-strategy routing with a non-uniform budget allocation algorithm to co-optimize strategy selection and resource distribution under a fixed total budget. Experiments on LLaMA-3.1-8B and Qwen3-8B demonstrate that, with an average cache budget of 512 tokens, PolyKV closes 54.5% and 25.7% of the LongBench performance gap between the strongest baseline and FullKV, respectively, and consistently outperforms all baselines by 1.7%–6.4% across a wide range of cache budgets.
This study addresses the substantial memory pressure and performance trade-offs in large language model (LLM) inference caused by KV cache growth with increasing context length and concurrent requests. The authors systematically evaluate three state-of-the-art KV cache management frameworks—vLLM, InfiniGen, and H2O—across diverse workloads, model scales, and sparsity conditions, measuring their latency, throughput, and memory efficiency. By integrating key techniques such as tensor offloading, token eviction heuristics, and speculative scheduling, the work provides the first comprehensive characterization of the operational boundaries of these strategies under varied deployment scenarios. It identifies optimal configurations under specific memory and performance constraints, offering empirical insights and practical guidance for designing efficient LLM inference systems.