π€ AI Summary
This study addresses the KV cache bottleneck in long-context inference for large language models and the performance degradation caused by existing methodsβ reliance on proxy signals. We propose a training-free KV cache compression framework that introduces, for the first time, the principle of preserving model prediction consistency as its core criterion. Specifically, critical cache entries are selected by evaluating the impact of candidate removals on the output distribution. Furthermore, we design a pre-eviction statistic reuse technique to eliminate the overhead of multiple forward passes. Experimental results demonstrate that, under identical cache budgets, the proposed framework significantly improves downstream task quality, with particularly pronounced advantages in aggressive compression scenarios, while simultaneously achieving end-to-end inference acceleration.
π Abstract
KV caches are a major bottleneck in long-context inference and long-form generation with large language models. Existing training-free eviction policies largely rely on proxy importance signals, such as attention mass, to decide which past tokens to retain. We argue that cache compression should instead preserve the predictive behavior of the full-cache model, retaining entries whose removal would substantially change the model's output distribution. We propose Behavior-Preserving KV Cache Compression, a training-free framework that scores candidate evictions by estimating the compressed-cache logits induced by their removal and evaluating the resulting KL to the full-cache next-token distribution. Using pre-eviction forward statistics, the method avoids running separate masked forward passes for each candidate. Across diverse architectures and both prefill-time and generation-time compression, our method delivers substantial gains in downstream task quality over lightweight attention-based heuristics at matched retained-KV budgets, with the largest gains under aggressive compression. It achieves these gains with additional compression-time computation while retaining an end-to-end speedup over full-cache inference in our evaluated settings.