🤖 AI Summary
This study addresses the interaction among position-agnostic missing blocks in cross-request KV caches by proposing a query-guided selective cache repair method. Moving beyond the constraints of fixed full-causal prefixes, this work formulates cache repair as a joint assignment problem. By leveraging sparse contextual attention and a query-guidance mechanism, it jointly optimizes layer-specific repair objectives with recomputed contexts. Experimental results demonstrate that the proposed approach outperforms ProphetKV on benchmarks such as LongBench, significantly improving the trade-off between generation quality and time-to-first-token latency.
📝 Abstract
Cross-request KV caching reduces the prefill cost of Retrieval-Augmented Generation (RAG), but conventional prefix caching severely limits cache reuse across requests. Position-Independent Caching (PIC) removes this constraint by reusing independent chunks, but their KV states miss cross-chunk interactions. Existing methods selectively recompute token states to recover these missing interactions, but primarily allocate the recomputation budget to selecting which states to recompute, while fixing the recomputation context to the full causal prefix. We introduce RelaxKV, which formulates selective cache repair as a joint allocation problem over repair targets and recomputation context. Guided by the user query, RelaxKV identifies layer-specific repair targets and restricts their recomputation to a query-relevant context, reducing attention computation. Across four decoder models, RelaxKV at a 15% anchor ratio improves aggregate LongBench performance over ProphetKV on all models. On Qwen3-14B, RelaxKV provides a stronger quality-TTFT trade-off than ProphetKV across a 5%-30% anchor-ratio sweep, and achieves the best selective results on RULER-MV and LV-Eval at 16K and 32K context lengths. Controlled ablations further demonstrate the importance of recomputation context selection.