build memory-managed attention

Designs and implements attention architectures that manage editable memory across long contexts, combining a fast recurrent or sparse backbone with explicit request-local memory slots and mechanisms for controlled writes, overwrites, and protections. Builds routing and query-time fallback logic that directs queries to a sparse datastore or sparse fallback and integrates hybrid long-context attention with memory-managed long-context attention to support efficient, editable retrieval and update semantics.

buildmemory-managedattention

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge that long-context language models struggle to distinguish between lossy historical compression and reliable long-term memory, as conventional attention mechanisms lack explicit control over memory writing, overwriting, protection, and forgetting. To overcome this limitation, the authors propose a novel memory management architecture that decouples memory lifecycle control from sequence processing for the first time. The design integrates a recurrent or sparse backbone with editable local memory slots and a query-time sparse fallback mechanism. Experiments demonstrate that this hybrid approach substantially outperforms purely fixed-state or purely sparse baselines on both synthetic and natural language tasks. Notably, small models achieve 595/600 accuracy under strong supervision and 1079/1080 pointer accuracy with frozen probes, confirming the efficacy of controllable memory slots and sparse fallbacks, while also highlighting open-domain memory selection as a critical remaining bottleneck.

editable memorylong-context attentionmemory lifecycle

This work addresses the inefficiency of traditional token-level context modeling in distinguishing between recollective, summarizing, and local information. The authors propose a novelty-driven memory mechanism that dynamically partitions context into three components: a content-addressable novelty cache for retrievable details, a recurrent state for compressed summaries, and a sliding window for recent local context. This architecture uniquely scales memory capacity with the amount of distinct information rather than raw token count, yielding an auditable and interpretable working memory structure. Integrating a Dirichlet-process-inspired novelty-gated attention mechanism, the system achieves full-attention performance in character-level control tasks with roughly half the attention cost and outperforms both full-attention and fixed-budget baselines on a thousand-event healthcare claims prediction task, while enabling human inspection of stored memory contents.

auditable memorycontext engineeringdistinct information

Traditional recurrent models struggle to effectively retain early contextual information in long sequences due to their fixed-size hidden states. This work proposes a state anchoring mechanism that periodically caches recurrent states as an expandable memory and generates content-conditional anchor keys for each cached state, enabling efficient retrieval through causal attention. The approach significantly enhances long-range memory capacity while maintaining computational efficiency. Experimental results demonstrate that the proposed method outperforms various linear attention variants across commonsense reasoning, LongBench, and context retrieval tasks, effectively improving the long-context modeling capabilities of recurrent architectures.

historical information retentionlong-context retrievalmemory recall

Current research on memory mechanisms in large language models remains highly fragmented and lacks a unified theoretical framework. This work proposes an architecture-centered taxonomy that systematically models memory along three orthogonal dimensions: representation, update dynamics, and persistence, formally characterizing core processes such as writing, routing, state transition, and integration. By constructing the first three-dimensional unified framework that integrates implicit/explicit, offline/online, and short-term/long-term memory, it clarifies the boundary between computationally coupled memory and independently addressable memory. Through a systematic literature review, architectural analysis, and multidimensional evaluation, the study synthesizes techniques including attention mechanisms, recurrent states, parameter-efficient fine-tuning, and scalable retrieval-augmented storage, thereby establishing a coherent paradigm for memory modeling and providing a theoretical foundation and design principles for future scalable and adaptive large language models.

architectural paradigmsfragmentationlarge language models

To address memory inefficiency and internal fragmentation in traditional KV caching for long-context reasoning in large language models (LLMs), this work proposes a fused attention mechanism integrating PagedAttention with PyTorch FlexAttention. Our novel design enables dynamic aggregation of disjoint KV blocks and supports non-contiguous memory layouts, eliminating fragmentation inherent in monolithic cache structures and enabling low-overhead long-sequence inference. Leveraging a custom CUDA fused kernel and integrated within the IBM Foundation Model Stack (FMS), experimental evaluation on an NVIDIA L4 GPU demonstrates near-linear latency scaling—approximately 2× increase—for sequences of 128–2048 tokens, while peak memory consumption remains nearly constant. A marginal increase in memory usage emerges only beyond 2048 tokens, attributable to power-of-two paging granularity.

Addresses memory inefficiencies in long-context LLM inferenceImproves linear latency scaling with global KV cachingIntegrates PagedAttention with FlexAttention to reduce fragmentation

Latest Papers

What's happening recently
View more

This work addresses the challenge of knowledge editing, which requires updating specific factual information while preserving unrelated yet semantically proximate knowledge. The authors propose a dual-adapter routing mechanism that employs a relevance router—supporting either lexical or BGE embeddings—to determine whether an input query matches stored edited knowledge. When a match is detected, an edit adapter is activated to prioritize the updated fact; otherwise, a locality adapter maintains the model’s original behavior. By decoupling the decisions of “when to write” and “when to suppress,” this approach enables more precise control over knowledge modifications. Implemented with parameter-efficient LoRA adapters, the method achieves state-of-the-art performance across three thousand-example benchmarks—CF, zsRE, and mQUAKE—on both Llama-3.1-8B-Instruct and Qwen3-8B, attaining a peak accuracy of 0.9922.

edit suppressionknowledge editinglocality preservation

This work addresses the limitations of sequence modeling approaches: state space models are constrained by fixed state dimensions, while attention mechanisms, despite their long-range memory capabilities, suffer from quadratic computational complexity and linearly growing cache overhead. To overcome these issues, the authors propose a sparse memory mechanism based on a learnable Dirichlet process cache that allocates memory slots only when inputs exhibit sufficient novelty, thereby scaling cache size with the number of distinct items rather than total token count. The method integrates DP-means clustering with a learnable novelty threshold gating mechanism, enabling end-to-end training. Experiments demonstrate that on redundant associative recall tasks, the approach matches full attention performance with substantially smaller cache sizes and outperforms fixed-budget methods in the memory-recall trade-off, validating the efficacy of “distinct-items” caching.

associative recallattention mechanismdistinct items

This work systematically investigates the interplay between efficient attention mechanisms—such as sliding window attention—and full attention in modern hybrid architectures, where the functional roles of these components remain poorly understood. Through extensive experiments, mechanistic analysis, and architectural ablation studies, the authors demonstrate that long-range information retrieval is predominantly handled by full attention layers, while efficient attention modules significantly shape the optimization trajectory. Building on these insights, they propose applying NoPE (No Positional Encoding) exclusively to full attention layers, which substantially enhances performance on long-context tasks with minimal degradation on short-context benchmarks. The study further uncovers a “large-window inertia” phenomenon, empirically validating that small-window sliding attention paired with NoPE-equipped full attention achieves superior efficiency and effectiveness.

efficient attentionfull attentionhybrid architectures

Hot Scholars

KD

Kishor Datta Gupta

Assistant Professor of Computer Science, Clark Atlanta University | Senior Member, IEEE
Physics Guided Neural NetworkPhysics Informed Machine LearningContext-Aware Machine learning
AR

Ahmed Rafi Hasan

ML Engineer @Pivotly | Graduate , United International University
Reinforcement LearningMachine LearningComputer VisionMultimodal Machine Learning
CZ

Chen Zhang

Shanghai Jiao Tong University
Power electronics systems stability
JZ

Jieru Zhao

Associate Professor, Shanghai Jiao Tong University
Hardware-software co-designAI acceleration and systemCompilerFPGA
SW

Shuo Wang

Institute of Automation, Chinese Academy of Sciences
RoboticsIntelligent RobotBiomimetic RobotMulti-Robot Systems