AvoKV-E: Payload-Aware KV Cache Eviction for Long Reasoning

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing KV cache eviction strategies in long reasoning, which overlook the impact of value payloads and often cause premature invalidation of newly generated states. To this end, it proposes a training-free eviction strategy that introduces value-payload awareness and a delayed observation mechanism into KV cache management for the first time. By deferring the eviction eligibility of recent states and heuristically ranking candidates based on normalized read pressure, key redundancy, and value payload potential, the method transcends the constraints of conventional routing-based perspectives. Under strict cache budgets, the proposed approach significantly outperforms existing baselines, demonstrating that preserving critical value payloads is essential for maintaining the integrity of long reasoning trajectories.
📝 Abstract
Long-output reasoning shifts the KV-cache bottleneck from the fixed prompt to the generated trace. Existing reasoning-cache eviction methods largely treat cached entries as routing objects, estimating whether an old key will still be read, will recur, or can be replaced. This routing-only view overlooks two effects: low-attention entries can carry large value payloads whose removal changes future predictions, and newly generated states can appear stale before later queries have had a chance to read them. We introduce AvoKV-E, a training-free eviction policy that first delays eligibility for recent states and then ranks eligible entries using candidate-normalized read pressure, key redundancy, and value-payload potential. According to empirical evaluation across different models and datasets, AvoKV-E matches or exceeds redundancy-aware, recurrence-based, and thought-adaptive eviction baselines at matched active-KV budgets, with its largest gains in the tightest-cache regime. Component and counterfactual analyses further connect these gains to delayed observation, payload-aware scoring, redundancy, and scale-robust normalization. Together, the results show that long-reasoning KV eviction should preserve not only keys that are likely to be read, but also the value payloads that sustain the reasoning trajectory.
Problem

Research questions and friction points this paper is trying to address.

KV cache eviction
long reasoning
value payload
cache management
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV Cache Eviction
Long Reasoning
Payload-Aware
Training-Free
Value-Payload Potential
🔎 Similar Papers
No similar papers found.