A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive memory overhead of KV caches in large language models, which severely constrains efficient deployment, alongside the limited expressiveness of existing decay mechanisms. To overcome these limitations, this work proposes a general attention framework that incorporates a trainable composite decay function. By fully preserving the representational capacity of RoPE and Softmax, the method achieves end-to-end adaptive KV cache pruning through a unified complementary decay mechanism. This approach transcends conventional bottlenecks, attaining a tenfold cache compression ratio while simultaneously improving model performance on standard tasks. In long-context scenarios, it further scales to achieve up to 25× compression. Consequently, the proposed method substantially optimizes both the inference and deployment efficiency of large language models.
📝 Abstract
The large KV-cache size of modern LLMs creates a barrier to efficient deployment. Recent work has explored replacing attention layers' RoPE positional embeddings with alternative decay-based mechanisms, which can then be used to prune KV-cache during inference. However, these decay functions have limited expressivity, and in practice devolve into sliding-window-like eviction patterns. In this work, we propose a unifying framework for complementary and novel decay mechanisms, capturing complex key statistics and interactions while preserving expressive RoPE embeddings and Softmax attention. The resulting Universal Attention is a highly expressive and end-to-end trainable architecture, whose composite decay mechanism acts as a natural, $\textit{adaptive}$ pruning criterion, removing tokens that contribute least to attention computation. Experimentally, Universal Attention achieves state-of-the-art $10\times$ compression on natural language and synthetic task data, while $\textit{improving}$ downstream performance compared to both state-of-the-art baselines and unpruned oracles. It further demonstrates superior long-context generalization with unprecedented $25\times$ compression at length 16k.
Problem

Research questions and friction points this paper is trying to address.

KV-cache compression
Large Language Models
efficient deployment
positional embeddings
attention mechanism
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV-Cache Compression
Universal Attention
Self-Pruning Transformer
Decay Mechanisms
RoPE
🔎 Similar Papers
No similar papers found.