Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration

πŸ“… 2026-10-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the trade-off between KV cache quantization and linear attention in Transformers regarding storage and computational costs, highlighting the absence of a unifying mechanism. To bridge this gap, we propose RAM-Net, which unifies per-KV compression and multi-KV aggregation through soft allocation over a discrete address space. Theoretically, we demonstrate that soft address allocation generalizes hard quantization matching and supports recurrent state updates, thereby establishing a principled weight transfer pathway from pretrained Transformers to RAM-Net. Empirically, across nine pretrained models, RAM-Net recovers 87.1% of the teacher model’s average accuracy gains with a fine-tuning budget of merely 500M tokens, achieving highly efficient attention.
πŸ“ Abstract
KV-cache quantization and linear attention are two representative approaches to tackling the storage and computational costs of Transformers. KV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contributions into a fixed-size continuous state but can introduce interference. This contrast raises the question of whether per-KV compression and multi-KV aggregation can be bridged within a single mechanism for efficient attention. We identify RAM-Net as such a bridge through soft assignments over a discrete address space. These assignments determine recurrent updates to the continuous slot state associated with each address. Under a restricted RAM-Net construction, we prove that soft address assignments extend hard quantized matching to a separable read-write overlap that locally approximates full-attention similarity and supports recurrent aggregation. These connections further enable Transformer-to-RAM-Net weight migration through a new path based on a soft-quantized intermediate construction. Across nine pretrained Transformer models from 0.3B to 7B parameters, RAM-Net recovers an average of 87.1% of the teachers' accuracy gains over random guessing across six commonsense and knowledge tasks using only a 500M-token budget per model.
Problem

Research questions and friction points this paper is trying to address.

KV-cache quantization
linear attention
efficient attention
weight migration
Transformer
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV-cache quantization
linear attention
RAM-Net
weight migration
soft assignment
πŸ”Ž Similar Papers