Budgeted Cache Repair for Cross-Context KV-Cache Reuse

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the accuracy degradation and absence of principled decision rules arising from cross-context KV cache reuse. We propose a budgeted cache repair method that reveals and quantifies the hidden costs of reuse, establishing the single row as the optimal decision granularity. Specifically, we introduce a draft attention mechanism to precisely identify corrupted cache rows, enabling fine-grained repair through attention-weight ranking and fixed-budget recomputation. Experimental results on the GSM8K benchmark demonstrate that our approach restores performance to dense prefill accuracy, significantly outperforming existing baselines. By effectively reconciling cache reuse efficiency with inference accuracy, this work provides a practical framework for optimizing KV cache management in large language models without compromising generation quality.
πŸ“ Abstract
Cross-context KV-cache reuse predicts a shared segment's keys and values under a new prefix instead of recomputing them, and has been reported to do so without quality loss. We find otherwise, and identify two problems. (1) A hidden cost: on MMLU and GSM8K, reuse costs substantial accuracy. (2) A decision at the wrong unit: no rule for deciding whether to reuse a cache removes that cost. What does help is choosing which parts of the cache to recompute, and the value of choosing well falls as the unit of choice grows: informed selection removes 49.5% of the cache error beyond chance at single rows (one token's keys and values), 10.6% at 64-token chunks, and nothing at the level of whole calls. Budgeted Cache Repair (BCR) acts at the unit where selection still pays. It drafts two tokens from the assembled cache, ranks cache rows by the attention those tokens pay them, and recomputes a fixed budget of rows exactly, in one of three layouts. The cost is paid rather than predicted away, and the draft that fails as a gate succeeds as a selector. BCR restores GSM8K to dense-prefill accuracy while still serving most calls from cache, and its best layout outperforms every reuse baseline's mean in the reference grid. The draft also beats a coin-flip selector at the same budget - a control prior evaluations lack.
Problem

Research questions and friction points this paper is trying to address.

KV-cache reuse
cross-context
cache repair
accuracy degradation
granularity
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV-Cache Reuse
Budgeted Cache Repair
Draft-and-Select
Attention-based Ranking
Cross-Context Inference
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.