🤖 AI Summary
This study addresses the evaluation distortion in KV cache reuse for Retrieval-Augmented Generation (RAG), where existing benchmark datasets lack complex reuse dynamics and thus fail to accurately quantify precision degradation. To overcome this limitation, this work proposes an unambiguous evaluation methodology and develops Boxoffice, a tool that leverages programmatic data synthesis to generate challenging benchmark datasets exhibiting intricate reuse patterns. These datasets effectively expose the inflated performance artifacts inherent in current evaluations. By establishing a rigorous assessment framework, this research achieves, for the first time, an authentic and precise measurement of accuracy loss incurred during KV cache reuse within RAG scenarios, thereby providing a more reliable foundation for evaluating retrieval-augmented systems.
📝 Abstract
Position-independent KV cache reuse aims to reduce latency in retrieval-augmented generation by reusing chunk-level KV caches across prompts. We show that current evaluations of KV cache reuse techniques rely on measurements that fail to faithfully capture the loss of accuracy attributable to reuse, often artificially inflating the reported effectiveness. We also show that existing datasets do not exhibit the reuse dynamics needed to thoroughly evaluate such techniques. To address these issues, we propose an evaluation methodology that measures this accuracy loss without ambiguity and we introduce Boxoffice, a tool that programmatically generates evaluation datasets that exercise challenging KV cache reuse patterns.