🤖 AI Summary
This study addresses the prohibitive memory overhead of KV caches in long-context models and the difficulty existing compression methods face in balancing computational cost with accuracy. To this end, it proposes a query-agnostic KV cache compression approach that introduces lightweight proxy scoring coupled with a selective recomputation mechanism. Specifically, an efficient filtering algorithm first identifies critical tokens, after which a subset undergoes localized score recomputation, enabling high-fidelity compression without requiring access to the full context. Experimental evaluations on benchmarks such as RULER demonstrate that the proposed method yields performance improvements exceeding 40% under extremely constrained memory budgets. Furthermore, it substantially reduces both runtime overhead and peak memory consumption, offering a practical solution for deploying long-context large language models efficiently.
📝 Abstract
The memory footprint of the key-value (KV) cache constrains the practical use of long-context models, and it dominates cost when one prefilled context must later serve many different queries. In this reusable setting, query-agnostic compression trades cost against quality: lightweight estimators are cheap but less accurate, whereas full-context reconstruction scoring is more accurate yet reprocesses the entire prompt. We introduce KV$^2$, a query-agnostic KV-cache compression method based on selective reconstruction. KV$^2$ first uses a lightweight proxy scorer to identify informative in-context tokens, then reprocesses only this subset to compute final eviction scores. On RULER, Needle-in-a-Haystack, and LongBench, KV$^2$'s margin over baselines widens as the budget tightens: on RULER 16K at a 2% KV-cache budget it improves the average score over the next-best baseline by more than 40 percentage points, and on LongBench it attains the highest average across 2%-10% budgets at lower compression-stage runtime and peak memory than full-context reconstruction. Reusable KV-cache compression thus does not require reprocessing the full context. Our code is available at https://anonymous.4open.science/r/KVsquared-0B97.