RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the interaction among position-agnostic missing blocks in cross-request KV caches by proposing a query-guided selective cache repair method. Moving beyond the constraints of fixed full-causal prefixes, this work formulates cache repair as a joint assignment problem. By leveraging sparse contextual attention and a query-guidance mechanism, it jointly optimizes layer-specific repair objectives with recomputed contexts. Experimental results demonstrate that the proposed approach outperforms ProphetKV on benchmarks such as LongBench, significantly improving the trade-off between generation quality and time-to-first-token latency.
📝 Abstract
Cross-request KV caching reduces the prefill cost of Retrieval-Augmented Generation (RAG), but conventional prefix caching severely limits cache reuse across requests. Position-Independent Caching (PIC) removes this constraint by reusing independent chunks, but their KV states miss cross-chunk interactions. Existing methods selectively recompute token states to recover these missing interactions, but primarily allocate the recomputation budget to selecting which states to recompute, while fixing the recomputation context to the full causal prefix. We introduce RelaxKV, which formulates selective cache repair as a joint allocation problem over repair targets and recomputation context. Guided by the user query, RelaxKV identifies layer-specific repair targets and restricts their recomputation to a query-relevant context, reducing attention computation. Across four decoder models, RelaxKV at a 15% anchor ratio improves aggregate LongBench performance over ProphetKV on all models. On Qwen3-14B, RelaxKV provides a stronger quality-TTFT trade-off than ProphetKV across a 5%-30% anchor-ratio sweep, and achieves the best selective results on RULER-MV and LV-Eval at 16K and 32K context lengths. Controlled ablations further demonstrate the importance of recomputation context selection.
Problem

Research questions and friction points this paper is trying to address.

KV cache reuse
Retrieval-Augmented Generation
Position-Independent Caching
selective recomputation
cross-chunk interactions
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV Cache Reuse
Sparse Context Attention
Selective Cache Repair
Retrieval-Augmented Generation
Position-Independent Caching
🔎 Similar Papers
R
Ruoling Qi
Shanghai Jiao Tong University; Institute of Artificial Intelligence, China Telecom (TeleAI)
Y
Yirui Liu
Institute of Artificial Intelligence, China Telecom (TeleAI)
X
Xuaner Wu
Institute of Artificial Intelligence, China Telecom (TeleAI)
Y
Yuxin Jin
Institute of Artificial Intelligence, China Telecom (TeleAI)
Jian Chen
Jian Chen
Associate Professor of Computer Science, The Ohio State University
Visualization3D interactionVirtual reality
Jiayu Qin
Jiayu Qin
University at Buffalo
machine learning
Yin Chen
Yin Chen
Lecturer in Mathematics at University of Saskatchewan
Invariant theoryLie theoryCommutative algebraApplied algebraic geometry
J
Jiawei Shao
Institute of Artificial Intelligence, China Telecom (TeleAI)