🤖 AI Summary
This work addresses the high memory cost of KV caching in large language model inference, where existing eviction strategies often rely on heuristics lacking theoretical grounding. The authors introduce, for the first time, the information bottleneck principle into KV cache management by constructing a linear Gaussian attention proxy model. This framework yields a mutual information–based objective function quantifying information capacity, leading to CapKV—an information-aware eviction method that directly optimizes the amount of useful information retained in the cache. The analysis reveals that several existing strategies are implicit approximations of this principle. Experiments across multiple models and long-context benchmarks demonstrate that CapKV consistently outperforms prior methods, achieving a superior trade-off between memory efficiency and generation fidelity.
📝 Abstract
Key-value (KV) caching is essential for large language model inference, yet its memory overhead poses a critical bottleneck for long-context generation. Existing eviction policies predominantly rely on empirical heuristics, lacking a rigorous theoretical foundation. This work rethinks KV cache eviction through the lens of the Information Bottleneck principle. Under a linear-Gaussian surrogate of attention, we derive a closed-form mutual information objective that characterizes the effective information capacity of a retained KV cache subset. This formulation reveals that a wide range of existing eviction strategies can be interpreted as different approximations of the same capacity-maximization principle. Guided by this insight, we introduce CapKV, a capacity-aware eviction method that directly targets information preservation via a log-determinant approximation using statistical leverage scores. This approach replaces heuristic selection with a theoretically grounded mechanism that preserves the maximum predictive signal. Extensive experiments across multiple models and long-context benchmarks show that CapKV consistently outperforms prior methods, achieving a better trade-off between memory efficiency and generational fidelity.