Behavior-Preserving KV Cache Compression

πŸ“… 2026-10-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the KV cache bottleneck in long-context inference for large language models and the performance degradation caused by existing methods’ reliance on proxy signals. We propose a training-free KV cache compression framework that introduces, for the first time, the principle of preserving model prediction consistency as its core criterion. Specifically, critical cache entries are selected by evaluating the impact of candidate removals on the output distribution. Furthermore, we design a pre-eviction statistic reuse technique to eliminate the overhead of multiple forward passes. Experimental results demonstrate that, under identical cache budgets, the proposed framework significantly improves downstream task quality, with particularly pronounced advantages in aggressive compression scenarios, while simultaneously achieving end-to-end inference acceleration.
πŸ“ Abstract
KV caches are a major bottleneck in long-context inference and long-form generation with large language models. Existing training-free eviction policies largely rely on proxy importance signals, such as attention mass, to decide which past tokens to retain. We argue that cache compression should instead preserve the predictive behavior of the full-cache model, retaining entries whose removal would substantially change the model's output distribution. We propose Behavior-Preserving KV Cache Compression, a training-free framework that scores candidate evictions by estimating the compressed-cache logits induced by their removal and evaluating the resulting KL to the full-cache next-token distribution. Using pre-eviction forward statistics, the method avoids running separate masked forward passes for each candidate. Across diverse architectures and both prefill-time and generation-time compression, our method delivers substantial gains in downstream task quality over lightweight attention-based heuristics at matched retained-KV budgets, with the largest gains under aggressive compression. It achieves these gains with additional compression-time computation while retaining an end-to-end speedup over full-cache inference in our evaluated settings.
Problem

Research questions and friction points this paper is trying to address.

KV Cache Compression
Large Language Models
Long-context Inference
Behavior Preservation
Training-free Eviction
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV Cache Compression
Behavior-Preserving
Training-Free
KL Divergence
Long-Context Inference
πŸ”Ž Similar Papers
No similar papers found.
D
Doo Hwan Hwang
Kim Jaechul Graduate School of AI, KAIST, Daejeon, Republic of Korea
J
Junyoung Jang
Kim Jaechul Graduate School of AI, KAIST, Daejeon, Republic of Korea
J
Junho Na
Kim Jaechul Graduate School of AI, KAIST, Daejeon, Republic of Korea
H
Hosung Lim
KT Corporation, Seoul, Republic of Korea
Kee-Eung Kim
Kee-Eung Kim
KAIST
Machine LearningReinforcement LearningDialogue Systems