🤖 AI Summary
This study addresses the accuracy degradation in KV cache retrieval for long-context reasoning caused by neglecting the impact of approximation errors on model outputs. To this end, we propose an output-aware page encoding method that pioneers a training-free retrieval mechanism directly modeling how approximation errors affect weighted attention outputs. By leveraging local page-based representations with joint key-value structures and spectral compression techniques, our approach optimizes sparse attention to achieve efficient and precise recall without requiring additional statistical information. Experimental results demonstrate that the proposed method significantly outperforms key-only baselines on long-context reasoning and understanding benchmarks. Furthermore, it achieves performance comparable to mainstream compression methods while maintaining minimal decoding overhead.
📝 Abstract
Long context inference with large language models becomes increasingly expensive as attention must operate over an ever growing KV cache. Page sparse attention reduces this cost by representing each KV page compactly and retrieving only a subset for each query. Existing retrieval methods are designed to estimate attention scores or page relevance, but their objectives do not directly account for how approximation errors affect the resulting value weighted attention output. We introduce \method{}, an output aware page encoding derived from the joint structure of keys and values while preserving the key information needed for accurate retrieval. \method{} is training free and requires no additional value dependent statistics at inference time. Once constructed, its stored representation has the same size and decode time scoring cost as a key only spectral representation. Across long reasoning, long context understanding, and long generation benchmarks, \method{} consistently improves over the key only spectral baseline and performs competitively with recent KV cache compression and retrieval methods. On long reasoning benchmarks, it achieves strong avg@\(k\) performance across model benchmark pairs, while matching or surpassing leading baselines on several long context understanding and generation settings with modest decoding overhead. Code is available at \url{https://github.com/Ashkan13776/oval-kv}.