🤖 AI Summary
This study addresses the degradation in generation quality within retrieval-augmented generation (RAG) caused by the loss of cross-chunk context when KV caches are computed and concatenated independently. To overcome this, we propose a novel, lightweight KV cache repair architecture based on residual learning. The network learns the discrepancy between independently and jointly computed caches by fusing compressed KV features with token embeddings, and employs a hybrid attention mechanism—bidirectional intra-chunk and forward inter-chunk—to predict and correct residuals without recomputing the target LLM. This approach supports both model-specific training and general corpus reuse, substantially reducing online overhead. Across 11 out of 12 evaluated configurations, our method achieves the Pareto frontier of quality versus latency, accelerating time-to-first-token generation by 1.69–4.61× while improving F1 scores by 2.1–26.1 percentage points.
📝 Abstract
Multi-document retrieval-augmented generation (RAG) requires a language model to process multiple retrieved text chunks before answering a question. Precomputing each chunk's KV cache independently and concatenating the caches when the chunks are retrieved can accelerate this step. However, the assembled cache lacks cross-chunk attention information, reducing answer quality. Selective recomputation methods recover the missing cross-chunk context by rerunning the target LLM on selected tokens, incurring substantial online computation. We introduce CacheRepair, a lightweight network that learns the difference between independently computed KV caches and those produced by processing the chunks together. The network combines compressed KV features with token embeddings and uses attention that is bidirectional within each chunk and flows from earlier to later chunks. Each repair block receives the compressed cache features, and the predicted residual is added to every document token's cache. Each repair network is trained for a specific frozen target LLM on a generic retrieval corpus and reused across downstream datasets. Our analysis shows that repair reduces KV errors both near chunk boundaries and throughout chunk interiors. Evaluation across three target LLMs and four downstream datasets places CacheRepair on the measured answer-quality-latency Pareto frontier in eleven of twelve model-dataset combinations. Reported time to first token (TTFT) includes online cache transfer and repair. Across all twelve combinations, the largest repairers achieve 1.69-4.61$\times$ speedups in median TTFT over full prefill and improve mean F1 by 2.1-26.1 percentage points over direct cache reuse.