Memory-Managed Long-Context Attention: A Preliminary Study of Editable Request-Local Memory

📅 2026-06-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that long-context language models struggle to distinguish between lossy historical compression and reliable long-term memory, as conventional attention mechanisms lack explicit control over memory writing, overwriting, protection, and forgetting. To overcome this limitation, the authors propose a novel memory management architecture that decouples memory lifecycle control from sequence processing for the first time. The design integrates a recurrent or sparse backbone with editable local memory slots and a query-time sparse fallback mechanism. Experiments demonstrate that this hybrid approach substantially outperforms purely fixed-state or purely sparse baselines on both synthetic and natural language tasks. Notably, small models achieve 595/600 accuracy under strong supervision and 1079/1080 pointer accuracy with frozen probes, confirming the efficacy of controllable memory slots and sparse fallbacks, while also highlighting open-domain memory selection as a critical remaining bottleneck.
📝 Abstract
Long-context language models often conflate two different goals: compressing history into an efficient state, and maintaining reliable long-term memory. Linear, recurrent, and sparse attention reduce the cost of processing long sequences, but they do not by themselves specify when a fact should be written, overwritten, protected from distractors, or discarded. We study memory-managed long-context attention, a research route that separates a fast recurrent or sparse backbone from explicit editable request-local memory slots and query-time sparse fallback. Across structured synthetic tasks, token/chunk/sequence bridges, generated natural language, and local frozen-model diagnostics, pure fixed-state or pure sparse methods fail some overwrite, version, anti-pollution, or no-write-signal cases, while a hybrid covers both routes. A small 2,097,152-token mechanism stress test reaches 50/50 pooled accuracy with 2-132 active chunks. A 2.74M-parameter minimal causal event-token model reaches 595/600 with lite write supervision, supporting proof of trainability rather than scale. A six-family frozen-hidden-state bridge reaches 1079/1080 controlled pointer accuracy, but it uses generator-provided integer key IDs and separately encoded canonical key strings; it is an oracle-metadata probe, not open-text entity resolution. Local non-leaderboard RULER 4K diagnostics remain close to full context, whereas a 33-record LongBench v1 16K subset shows that naive lexical selection is not general. The evidence separates three claims: controlled slot lifecycle is feasible, sparse fallback is needed when writes lack future-query signals, and learned open-domain selection remains the main architectural bottleneck. We do not claim a final generative architecture, global slot-trajectory convergence, or systems superiority.
Problem

Research questions and friction points this paper is trying to address.

long-context attention
memory management
editable memory
sparse attention
memory lifecycle
Innovation

Methods, ideas, or system contributions that make the work stand out.

memory-managed attention
editable request-local memory
sparse fallback
slot lifecycle
long-context modeling