Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective

📅 2026-04-28
📈 Citations: 0
Influential: 0
📄 PDF

career value

188K/year
🤖 AI Summary
This work addresses the high memory cost of KV caching in large language model inference, where existing eviction strategies often rely on heuristics lacking theoretical grounding. The authors introduce, for the first time, the information bottleneck principle into KV cache management by constructing a linear Gaussian attention proxy model. This framework yields a mutual information–based objective function quantifying information capacity, leading to CapKV—an information-aware eviction method that directly optimizes the amount of useful information retained in the cache. The analysis reveals that several existing strategies are implicit approximations of this principle. Experiments across multiple models and long-context benchmarks demonstrate that CapKV consistently outperforms prior methods, achieving a superior trade-off between memory efficiency and generation fidelity.
📝 Abstract
Key-value (KV) caching is essential for large language model inference, yet its memory overhead poses a critical bottleneck for long-context generation. Existing eviction policies predominantly rely on empirical heuristics, lacking a rigorous theoretical foundation. This work rethinks KV cache eviction through the lens of the Information Bottleneck principle. Under a linear-Gaussian surrogate of attention, we derive a closed-form mutual information objective that characterizes the effective information capacity of a retained KV cache subset. This formulation reveals that a wide range of existing eviction strategies can be interpreted as different approximations of the same capacity-maximization principle. Guided by this insight, we introduce CapKV, a capacity-aware eviction method that directly targets information preservation via a log-determinant approximation using statistical leverage scores. This approach replaces heuristic selection with a theoretically grounded mechanism that preserves the maximum predictive signal. Extensive experiments across multiple models and long-context benchmarks show that CapKV consistently outperforms prior methods, achieving a better trade-off between memory efficiency and generational fidelity.
Problem

Research questions and friction points this paper is trying to address.

KV cache eviction
memory overhead
long-context generation
information bottleneck
theoretical foundation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Information Bottleneck
KV Cache Eviction
Mutual Information
Statistical Leverage Scores
Long-Context Generation