KV-Kaizen: Learning Context-Adaptive Cache Compression Choices

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the excessive memory consumption of KV caches during long-context inference in large language models, which severely constrains system throughput. To overcome this limitation, it proposes a learnable, context-aware selector that abandons monolithic eviction strategies in favor of dynamically generating layer-adaptive cache configurations. By integrating low-rank decomposition, quantization, and cross-layer sharing techniques, the method achieves localized and adaptive compression interventions across three dimensions: depth, precision, and rank. Evaluated on a 14B-parameter model, the proposed approach attains a 32-fold cache reduction without accuracy degradation, reaching the Pareto optimal frontier between accuracy and cache overhead. Ultimately, this work establishes a new paradigm for efficient long-context inference.
📝 Abstract
As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size. Recent work alleviates this bottleneck by discarding the least relevant tokens. Eviction introduces a tension, since a one-off decision to discard content may prove detrimental later. Instead, we focus on alternative choices that can lead to cache compression without evicting tokens. We achieve this by learning a selector that is able to produce, based on context, a per-layer cache configuration towards an overall compression budget. The selector operates along three axes: sharing one cache across layers (depth), caching at fewer bits (precision), or truncating the low-rank latent cache representations (rank). We call the resulting method KV-Kaizen, for the many small per-layer choices it compounds. We observe that these interventions taken independently and uniformly over all layers limit achievable compression because they degrade accuracy. Crucially, composing them locally and adaptively to the context can instead preserve accuracy while achieving large memory savings. At inference, the selector runs once, before pre-fill. In evaluations on instruction following and reasoning tasks, our selectors reach the Pareto frontier of accuracy against cache size, against learning-free and post-hoc baselines. On long-context tasks, KV-Kaizen improves on eviction and can be composed with it, reaching a 32x smaller decode-time cache on a 14B model while preserving accuracy. A 4x cache size reduction incurs no accuracy degradation from 7B parameters up, and a compressed model is more accurate than a smaller uncompressed one with the same cache size. Together, these findings support pre-training large models and compressing them only afterwards.
Problem

Research questions and friction points this paper is trying to address.

KV cache compression
large language models
long context
memory bottleneck
inference throughput
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV cache compression
context-adaptive selector
low-rank truncation
precision reduction
cross-layer sharing