🤖 AI Summary
This study addresses the excessive memory consumption of KV caches in long-context Transformers by proposing AttSVD. This method introduces a novel attention-guided, interpretable low-rank compression paradigm that performs online truncated singular value decomposition (SVD) based on prompt attention geometry, compressing the KV cache along the feature dimension. By incorporating an energy rule and a perception-based truncation mechanism alongside cumulative and streaming decoding strategies, AttSVD achieves adaptive compression while retaining all tokens, thereby providing interpretability insights at no additional cost. Experimental results demonstrate that AttSVD matches the performance of dense caching on benchmarks such as LongBench while reducing KV cache memory usage by up to 50%.
📝 Abstract
The key-value (KV) cache of autoregressive transformers grows linearly with context length and dominates memory at long context. Most training-free remedies evict low-importance tokens, an irreversible choice along the sequence axis. We instead keep every token and store it more cheaply along the "feature" axis. We therefore propose AttSVD, a new "interpretable" low-rank compression whose basis is derived from each prompt's own attention geometry: an online, per-prompt truncated SVD that keeps only the directions attention actually reads, cutting persistent per-head KV memory in proportion to the retained rank. We propose two decode-time caching strategies, accumulating and streaming, for short and long generation regimes. Furthermore, we propose two refinements that make compression adaptive. A per-matrix energy rule sizes the logit space and the attention mass independently. An attention-aware basis truncates only in the spaces attention actually reads, preserving both the attention logits and the attention output. The same factors also provide free, per-head interpretability insights into the effective rank and the geometry attention consumes. Across multiple models, on both an agentic benchmark and the full LongBench suite AttSVD stays on par with the dense cache while using up to 50% of the KV-cache memory.