More Than a Quick Glance: Overcoming the Greedy Bias in KV-Cache Compression

📅 2026-02-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing KV cache compression methods often sacrifice semantic recall due to their reliance on attention scores and sliding window mechanisms, and they face challenges in fair performance evaluation. This work proposes LASER-KV, a novel framework that introduces a block-level cumulative budget mechanism and a dynamic protection factor to decouple compression from sliding window interference. By integrating Exact-LSH for semantic-aware compression, LASER-KV achieves high recall without resorting to fixed-size greedy strategies, instead employing hierarchical selection for more rational cache budget allocation. Evaluated on the Babilong benchmark with a 128k context length, LASER-KV improves accuracy by up to 10% over current methods and effectively mitigates performance degradation of 15–30%.

Technology Category

Data Mining & Knowledge Management: Data CompressionMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Summarization

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Large language models for searchEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
While Large Language Models (LLMs) can theoretically support extensive context windows, their actual deployment is constrained by the linear growth of Key-Value (KV) cache memory. Prevailing compression strategies mitigate this through various pruning mechanisms, yet trade-off semantic recall for memory efficiency. In this work, we present LASER-KV (Layer Accumulated Selection with Exact-LSH Recall), a framework designed to test the limits of KV compression under a strict accumulative budgeting policy. We deviate from the standard fixed summary size approach by implementing a block-wise accumulation strategy governed by a protection divisor (n). This allows us to isolate the effects of compression from sliding window artifacts. Our experiments on the Babilong benchmark reveal performance degradation in previous compression methods by 15-30% on various long context tasks. LASER-KV maintains stable performance, achieving superior accuracies by a margin of upto 10% at 128k. These findings challenge the prevailing assumption that attention scores alone are a sufficient proxy for token utility.
Problem

Research questions and friction points this paper is trying to address.

KV-cache compression
memory efficiency
long context
semantic recall
attention mechanisms
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV-cache compression
accumulative budgeting
Exact-LSH Recall
long-context modeling
attention mechanism
🔎 Similar Papers
No similar papers found.
A
Aryan Sood
Indian Institute of Technology, Roorkee
T
Tanvi Sharma
Indian Institute of Technology, Roorkee
V
Vansh Agrawal
Indian Institute of Technology, Roorkee