When to Evict, Not What to Keep: Draft-Guided Eviction for Training-Free KV-Cache Compression

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of existing KV cache compression methods, where premature eviction of critical tokens causes attention compensation and selection mechanisms to fail. We propose Draft-Guided Eviction (DGE), which shifts the optimization focus from "what to retain" to "when to evict." By leveraging draft generation to introduce a trajectory anchoring effect, DGE effectively resolves the issue of missing early signals and enables deferred eviction. Notably, it supports training-free KV cache compression and head-level budget control without modifying scoring functions. Experimental results demonstrate that DGE outperforms existing baselines across most models, achieving a LongBench score of 44.2 and closely approaching full-cache performance.
πŸ“ Abstract
Training-free KV-cache compression methods such as SnapKV, H2O, and PyramidKV evict tokens at the end of prefill, aiming to preserve the attention mass that future queries are expected to use -optimizing what to keep. We show that this objective fails in two distinct ways. (1) Compensation: restoring the evicted attention mass can recover the attention-level target without recovering task quality. (2) Selection: covering more of the true decode-query mass can hurt quality when the recovered mass is fragmented rather than concentrated in coherent spans. These failures share a common cause: eviction occurs before the queries that determine the answer trajectory exist. We propose Draft-Guided Eviction (DGE), which defers eviction until after drafting the first k=2 answer tokens using the full cache - just one decode step beyond prefill. Because the draft is generated from the answer's own prefix, no cache entries are discarded before this trajectory signal becomes available. The per-head cache budget remains unchanged, and DGE can be applied directly to SnapKV, PyramidKV, H2O, and StreamingLLM without modifying their eviction scores. Unlike extra-pass methods, DGE changes when eviction occurs rather than what cache entries are selected. Extensive experiments demonstrate that DGE outperforms prior methods at every evaluated budget on five of six instruct-tuned backbones, achieving 44.2 on LongBench, nearly matching FullKV at 44.3. The timing-only control DGE-W achieves the same score, demonstrating that the gain comes from when eviction occurs rather than what is selected - an effect we term trajectory anchoring.
Problem

Research questions and friction points this paper is trying to address.

KV-cache compression
training-free
token eviction
attention mass
trajectory anchoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

Draft-Guided Eviction
KV-Cache Compression
Training-Free
Trajectory Anchoring
LLM Inference