How Linear Attention Remembers

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the memory interference and capacity bottlenecks inherent in the fixed-size recurrent states of linear attention mechanisms. Through analytical decomposition and causal intervention, we systematically investigate the writing, retention, and retrieval dynamics in models such as Gated Linear Attention (GLA), providing the first quantification of cross-fact causal coupling. Our findings reveal that interference stems primarily from overlapping subsequent writes rather than temporal decay, challenging conventional understandings of KV caching. Furthermore, we demonstrate that both recall and editing performance degrade under high cognitive load, and identify that full-attention layers predominantly govern real-time retrieval within hybrid architectures. Ultimately, this work delineates the memory boundaries and selective recall principles of linear attention, offering critical insights for advancing state-space models.
📝 Abstract
Linear attention replaces the growing key--value (KV) cache of standard attention with a fixed-size recurrent state, substantially reducing memory growth with context length. This efficiency, however, changes how past information is stored: many tokens must share and repeatedly update the same memory. We study how this recurrent state functions as a memory system. Using an analytical decomposition together with controlled causal interventions in pretrained GLA and GDN models, we trace how recalled information is written, retained, and later accessed. We find that fact-specific information enters recurrent memory through concentrated, content-dependent writes and is later accessed through concentrated query-time read pathways. Multiple facts can remain selectively accessible within the same state, yet their internal representations exhibit cross-fact causal coupling rather than independent KV-like storage. As memory load increases, both recall and targeted editability degrade, whereas elapsed context alone has a substantially smaller effect within the tested regime. Causal interventions on subsequent writes further show that interference is shaped by their overlap with existing memory. Finally, in hybrid architectures that combine recurrent layers with full attention, the runtime memory directly supporting recall shifts predominantly to the full-attention KV state. Together, these results reveal how fixed-size recurrent memory supports selective recall despite shared storage, while exposing the interference and capacity limits that distinguish it from token-addressable KV memory.
Problem

Research questions and friction points this paper is trying to address.

Linear Attention
Recurrent Memory
KV Cache
Memory Interference
Information Recall
Innovation

Methods, ideas, or system contributions that make the work stand out.

Linear Attention
Recurrent Memory
Causal Interventions
KV Cache
Memory Interference
🔎 Similar Papers
No similar papers found.