Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning

📅 2024-10-25
🏛️ International Conference on Learning Representations
📈 Citations: 46
✨ Influential: 6
📄 PDF
🤖 AI Summary
To address the quadratic memory overhead of KV caches in large language models (LLMs) with increasing context length, this paper proposes HeadKV-R2, a head-granularity dynamic compression method. HeadKV-R2 introduces the first attention-head-level KV cache compression paradigm, leveraging a context-aware importance scoring mechanism that jointly models retrieval and reasoning capabilities. Guided by question-answering tasks, it quantifies the contribution of each attention head and performs dynamic sub-sampling of KV tokens per head. Evaluated on LongBench and LooGLE benchmarks across Llama-3-8B and Mistral-7B, HeadKV-R2 significantly outperforms layer-wise compression baselines: it retains 97% of full-cache performance while preserving only 1.5% of the original KV cache, and achieves up to 12.7% absolute improvement when KV cache sizes are set to 64 or 128. The implementation is publicly available.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Data Mining & Knowledge Management: Data CompressionComputer Vision: Large Vision Models

Application Category

Search and Retrieval-Augmented AI: Large language models for searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Key-Value (KV) caching is a common technique to enhance the computational efficiency of Large Language Models (LLMs), but its memory overhead grows rapidly with input length. Prior work has shown that not all tokens are equally important for text generation, proposing layer-level KV cache compression to selectively retain key information. Recognizing the distinct roles of attention heads in generation, we propose HeadKV, a head-level KV cache compression method, and HeadKV-R2, which leverages a novel contextual reasoning ability estimation for compression. Our approach operates at the level of individual heads, estimating their importance for contextual QA tasks that require both retrieval and reasoning capabilities. Extensive experiments across diverse benchmarks (LongBench, LooGLE), model architectures (e.g., Llama-3-8B-Instruct, Mistral-7B-Instruct), and long-context abilities tests demonstrate that our head-level KV cache compression significantly outperforms strong baselines, particularly in low-resource settings (KV size = 64&128). Notably, our method retains just 1.5% of the KV cache while achieving 97% of the performance of the full KV cache on the contextual question answering benchmark.Codes are available at https://github.com/FYYFU/HeadKV
Problem

Research questions and friction points this paper is trying to address.

Compressing KV cache memory overhead in large language models
Selectively retaining important attention heads for generation tasks
Maintaining performance with minimal KV cache in retrieval-reasoning tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compresses KV cache at individual attention head level
Uses contextual reasoning ability estimation for compression
Retains minimal KV cache while maintaining high performance
🔎 Similar Papers
No similar papers found.