Score
Designs and implements attention architectures that manage editable memory across long contexts, combining a fast recurrent or sparse backbone with explicit request-local memory slots and mechanisms for controlled writes, overwrites, and protections. Builds routing and query-time fallback logic that directs queries to a sparse datastore or sparse fallback and integrates hybrid long-context attention with memory-managed long-context attention to support efficient, editable retrieval and update semantics.
The quadratic time and memory complexity of Transformer self-attention severely hinders efficient long-context modeling. This paper presents a systematic survey and reconstruction of efficient attention mechanisms for large language models, proposing the first unified taxonomy encompassing both linearization paradigms (e.g., kernel-based approximations and fast weight dynamics) and sparsification paradigms (e.g., fixed patterns, block-wise routing, and clustering-driven selection). It innovatively integrates algorithmic design with hardware-aware optimization, clarifying integration pathways for purely efficient attention and hybrid architectures in large-scale pretraining. Furthermore, it establishes a comprehensive reference framework spanning theoretical analysis, algorithmic implementation, and engineering deployment. The work delivers a systematic design paradigm and practical guidelines for scalable long-context language models.
This work addresses the challenge that long-context language models struggle to distinguish between lossy historical compression and reliable long-term memory, as conventional attention mechanisms lack explicit control over memory writing, overwriting, protection, and forgetting. To overcome this limitation, the authors propose a novel memory management architecture that decouples memory lifecycle control from sequence processing for the first time. The design integrates a recurrent or sparse backbone with editable local memory slots and a query-time sparse fallback mechanism. Experiments demonstrate that this hybrid approach substantially outperforms purely fixed-state or purely sparse baselines on both synthetic and natural language tasks. Notably, small models achieve 595/600 accuracy under strong supervision and 1079/1080 pointer accuracy with frozen probes, confirming the efficacy of controllable memory slots and sparse fallbacks, while also highlighting open-domain memory selection as a critical remaining bottleneck.
This work addresses the inefficiency of traditional token-level context modeling in distinguishing between recollective, summarizing, and local information. The authors propose a novelty-driven memory mechanism that dynamically partitions context into three components: a content-addressable novelty cache for retrievable details, a recurrent state for compressed summaries, and a sliding window for recent local context. This architecture uniquely scales memory capacity with the amount of distinct information rather than raw token count, yielding an auditable and interpretable working memory structure. Integrating a Dirichlet-process-inspired novelty-gated attention mechanism, the system achieves full-attention performance in character-level control tasks with roughly half the attention cost and outperforms both full-attention and fixed-budget baselines on a thousand-event healthcare claims prediction task, while enabling human inspection of stored memory contents.
Traditional recurrent models struggle to effectively retain early contextual information in long sequences due to their fixed-size hidden states. This work proposes a state anchoring mechanism that periodically caches recurrent states as an expandable memory and generates content-conditional anchor keys for each cached state, enabling efficient retrieval through causal attention. The approach significantly enhances long-range memory capacity while maintaining computational efficiency. Experimental results demonstrate that the proposed method outperforms various linear attention variants across commonsense reasoning, LongBench, and context retrieval tasks, effectively improving the long-context modeling capabilities of recurrent architectures.
Current research on memory mechanisms in large language models remains highly fragmented and lacks a unified theoretical framework. This work proposes an architecture-centered taxonomy that systematically models memory along three orthogonal dimensions: representation, update dynamics, and persistence, formally characterizing core processes such as writing, routing, state transition, and integration. By constructing the first three-dimensional unified framework that integrates implicit/explicit, offline/online, and short-term/long-term memory, it clarifies the boundary between computationally coupled memory and independently addressable memory. Through a systematic literature review, architectural analysis, and multidimensional evaluation, the study synthesizes techniques including attention mechanisms, recurrent states, parameter-efficient fine-tuning, and scalable retrieval-augmented storage, thereby establishing a coherent paradigm for memory modeling and providing a theoretical foundation and design principles for future scalable and adaptive large language models.
To address memory inefficiency and internal fragmentation in traditional KV caching for long-context reasoning in large language models (LLMs), this work proposes a fused attention mechanism integrating PagedAttention with PyTorch FlexAttention. Our novel design enables dynamic aggregation of disjoint KV blocks and supports non-contiguous memory layouts, eliminating fragmentation inherent in monolithic cache structures and enabling low-overhead long-sequence inference. Leveraging a custom CUDA fused kernel and integrated within the IBM Foundation Model Stack (FMS), experimental evaluation on an NVIDIA L4 GPU demonstrates near-linear latency scaling—approximately 2× increase—for sequences of 128–2048 tokens, while peak memory consumption remains nearly constant. A marginal increase in memory usage emerges only beyond 2048 tokens, attributable to power-of-two paging granularity.
This work addresses the challenge of knowledge editing, which requires updating specific factual information while preserving unrelated yet semantically proximate knowledge. The authors propose a dual-adapter routing mechanism that employs a relevance router—supporting either lexical or BGE embeddings—to determine whether an input query matches stored edited knowledge. When a match is detected, an edit adapter is activated to prioritize the updated fact; otherwise, a locality adapter maintains the model’s original behavior. By decoupling the decisions of “when to write” and “when to suppress,” this approach enables more precise control over knowledge modifications. Implemented with parameter-efficient LoRA adapters, the method achieves state-of-the-art performance across three thousand-example benchmarks—CF, zsRE, and mQUAKE—on both Llama-3.1-8B-Instruct and Qwen3-8B, attaining a peak accuracy of 0.9922.
This work addresses the limitations of sequence modeling approaches: state space models are constrained by fixed state dimensions, while attention mechanisms, despite their long-range memory capabilities, suffer from quadratic computational complexity and linearly growing cache overhead. To overcome these issues, the authors propose a sparse memory mechanism based on a learnable Dirichlet process cache that allocates memory slots only when inputs exhibit sufficient novelty, thereby scaling cache size with the number of distinct items rather than total token count. The method integrates DP-means clustering with a learnable novelty threshold gating mechanism, enabling end-to-end training. Experiments demonstrate that on redundant associative recall tasks, the approach matches full attention performance with substantially smaller cache sizes and outperforms fixed-budget methods in the memory-recall trade-off, validating the efficacy of “distinct-items” caching.
This work systematically investigates the interplay between efficient attention mechanisms—such as sliding window attention—and full attention in modern hybrid architectures, where the functional roles of these components remain poorly understood. Through extensive experiments, mechanistic analysis, and architectural ablation studies, the authors demonstrate that long-range information retrieval is predominantly handled by full attention layers, while efficient attention modules significantly shape the optimization trajectory. Building on these insights, they propose applying NoPE (No Positional Encoding) exclusively to full attention layers, which substantially enhances performance on long-context tasks with minimal degradation on short-context benchmarks. The study further uncovers a “large-window inertia” phenomenon, empirically validating that small-window sliding attention paired with NoPE-equipped full attention achieves superior efficiency and effectiveness.