design long-term memory

Design, build, or analyze mechanisms and system architectures that let algorithms and models store, retrieve, and update information over long time horizons or extended input contexts — this includes recurrent memory units (e.g., long short-term memory), external or episodic memory modules, indexing and retrieval layers, and session-level persistence. Work covers evaluating and engineering capacity and latency trade‑offs, forgetting and consistency semantics, compression and retrieval accuracy, memory management policies across training and inference, and integration of long‑context memory with model optimization and deployment.

designlong-termmemory

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$219K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Current research on memory mechanisms in large language model (LLM) agents is fragmented across operating systems engineering and cognitive science, lacking a unified evolutionary perspective. This work proposes a three-stage memory evolution framework—storage, reflection, and experience—that systematically integrates recent advances in the field and formally defines the core drivers and key capabilities of each stage, such as active exploration and cross-trajectory abstraction. By synthesizing theoretical insights from cognitive science and systems engineering through comprehensive review and framework-based modeling, this study establishes a unified evolutionary theory of memory for LLM agents. The resulting framework offers clear design principles and a developmental roadmap for next-generation agents, advancing memory systems from passive recording toward active experience generation.

continual learningevolutionary frameworkLLM agent

This work addresses the challenges faced by foundational agents in long-horizon, dynamic, and user-dependent environments—particularly context explosion and sustained information management—which necessitate efficient memory mechanisms to enhance practical utility. The paper proposes the first unified three-dimensional framework that integrates internal and external memory, five cognitive mechanisms, and a dual-agent/user-centric perspective, offering a systematic structuring of research on agent memory. Drawing on a comprehensive review of hundreds of studies published before 2026, the framework synthesizes insights from memory modeling, cognitive science, and agent architecture to clarify memory operation strategies, evaluation benchmarks, and learning methodologies. It not only delineates structured pathways in existing research but also identifies key open problems, thereby providing a theoretical foundation and directional guidance for the future design of intelligent agent memory systems.

context explosionfoundation agentslong-horizon environments

Must-Read Papers

Most classic and influential ideas
View more

Structured Memory Mechanisms for Stable Context Representation in Large Language Models

May 28, 2025
YX
Yue Xing
🏛️ University of Pennsylvania | University of Michigan | University of Saskatchewan | Fordham University | Northeastern University

To address semantic loss, drift, and context decay in large language models (LLMs) during long-context understanding, this paper proposes an explicit structured long-term memory architecture. Methodologically, it introduces a novel memory unit integrating gated writing, attention-driven reading, and a learnable dynamic forgetting function, coupled with a multi-objective training framework that jointly optimizes task performance and memory policy. Crucially, the approach incurs no additional inference latency. Experiments demonstrate substantial improvements in long-text generation coherence, multi-turn dialogue stability, and cross-paragraph reasoning accuracy. The memory mechanism is empirically validated for its effectiveness in semantic persistence and precise retrieval, exhibiting strong generalization across diverse long-context tasks. This work establishes a new paradigm for modeling extended contextual dependencies in LLMs.

Enhancing long-term context understanding in large language modelsImproving semantic retention and retrieval across paragraphs and dialoguesMitigating context loss and semantic drift in long-text tasks

Multiple Memory Systems for Enhancing the Long-term Memory of Agent

Aug 21, 2025
GZ
Gaoke Zhang
🏛️ Tianjin University | China Mobile Communication Group Tianjin Co., Ltd | Huizhi Xingyuan Information Technology Co., Ltd

To address the low fidelity of long-term memory and inefficient historical data utilization in large language model (LLM) agents, this paper proposes a cognitively inspired Multi-Memory System (MMS). MMS dynamically transforms short-term memory into structured long-term memory fragments and introduces a novel one-to-one correspondence mechanism between retrieval memory units and contextual memory units, enabling high-precision matching and efficient recall. The method integrates techniques including MemoryBank and A-MEM, and incorporates a hierarchical memory architecture grounded in cognitive theory to jointly optimize storage efficiency and semantic coherence. On the LoCoMo benchmark, MMS significantly outperforms three baseline approaches. Ablation studies validate the effectiveness of the memory unit design, and empirical results demonstrate robustness and practicality across varying memory capacities and storage overheads.

Efficiently processing vast historical data for retrievalEnhancing agent long-term memory quality from interactionsImproving recall performance and response quality issues

Current research on memory mechanisms in large language models remains highly fragmented and lacks a unified theoretical framework. This work proposes an architecture-centered taxonomy that systematically models memory along three orthogonal dimensions: representation, update dynamics, and persistence, formally characterizing core processes such as writing, routing, state transition, and integration. By constructing the first three-dimensional unified framework that integrates implicit/explicit, offline/online, and short-term/long-term memory, it clarifies the boundary between computationally coupled memory and independently addressable memory. Through a systematic literature review, architectural analysis, and multidimensional evaluation, the study synthesizes techniques including attention mechanisms, recurrent states, parameter-efficient fine-tuning, and scalable retrieval-augmented storage, thereby establishing a coherent paradigm for memory modeling and providing a theoretical foundation and design principles for future scalable and adaptive large language models.

architectural paradigmsfragmentationlarge language models

MemOS: A Memory OS for AI System

Jul 04, 2025
ZL
Zhiyu Li
🏛️ MemTensor Technology Co., Ltd. | Research Institute of China Telecom | Renmin University of China | Peking University

Large language models (LLMs) lack a unified memory management framework, hindering long-context reasoning, continual personalization, and knowledge consistency. Existing approaches—relying on static parameters, transient context windows, or stateless retrieval-augmented generation (RAG)—fail to support heterogeneous knowledge co-evolution across temporal scales and sources. Method: We propose MemOS, a memory operating system for LLMs, introducing the MemCube as a foundational memory abstraction that unifies explicit, activation-level, and parameter-level memory representations, enabling composition, migration, and fusion. MemOS employs hierarchical modeling integrating retrieval, activation control, and parameter updates to realize a schedulable, evolvable, lifecycle-aware memory management architecture. Results: Experiments demonstrate that MemOS significantly reduces training and inference overhead, improves knowledge update efficiency and context management capability, and provides a scalable, controllable memory infrastructure for continual learning and personalized modeling.

Current systems fail to manage heterogeneous knowledge efficiently.Lack of memory management in LLMs hinders long-context reasoning.Static parameters limit tracking user preferences over time.

Latest Papers

What's happening recently
View more

This work addresses the efficiency challenges of cross-session memory management in large language model agents during long-horizon tasks, a domain where existing systems lack systematic understanding of memory mechanisms. The study introduces, for the first time, a four-dimensional taxonomy for agent memory and a phase-aware performance profiling framework. Leveraging two benchmark suites, the authors empirically evaluate ten representative systems, uncovering cost distribution patterns across memory construction, retrieval, and generation phases. The analysis quantifies trade-offs in write and read path overheads across different architectures and distills ten system design principles encompassing construction scheduling, capability baselines, query amortization, freshness–latency balance, and cluster management.

Agent MemoryLong-Horizon TasksMemory Systems

This work addresses the cognitive degradation and escalating computational overhead in large language models during extended scientific collaboration, caused by context saturation. To overcome these limitations, we propose a dual-process memory architecture that decouples short-term episodic memory (fixed to the latest 10 messages) from long-term semantic knowledge (growing at approximately 3 tokens per message). The framework integrates domain-specific knowledge compression, dual-channel episodic-semantic memory, and cross-model verification, enabling robust handling of parameter contradictions, multi-hop reasoning across collaboration stages, and precise retention of technical facts. Evaluated across six mainstream large language models, our system maintains 70–85% accuracy over 15,000 messages with only 1–2 seconds of latency, reduces token consumption by 62%, and successfully manages over 14,000 scientific facts (125k tokens), substantially surpassing the capacity and efficiency limits of conventional full-context approaches.

context window saturationlarge language modelslong-horizon scientific agents

This work addresses the lack of effective long-term memory management in large language model agents during extended interactions. Inspired by human cognition, the authors propose a novel memory architecture that systematically integrates six key mechanisms: sleep-based consolidation, interference-driven forgetting, memory trace maturation, retrieval-induced reconsolidation, entity-centric knowledge graphs, and multi-cue hybrid retrieval. To prevent data leakage, they introduce an unsupervised synthetic calibration method for threshold setting. The framework further incorporates memory deduplication and compression, context budget control, and a streaming multi-level evaluation paradigm. Evaluated on the VSCode dataset, the approach achieves 97.2% memory retention accuracy while reducing storage overhead by 58%. On the LongMemEval benchmark, it matches baseline retrieval accuracy using only 200K tokens of context and improves S-tier preference recall by 13.3 percentage points.

LLM agentslong interaction horizonsmemory consolidation

This study addresses the lack of fine-grained, system-level evaluation in existing memory systems for large language model agents, which hinders robust assessment under cost constraints, architectural trade-offs, and dynamic knowledge updates. The work proposes a unified analytical framework by decomposing agent memory into four independently evaluable core modules: representation storage, extraction, retrieval routing, and maintenance. Through end-to-end and ablation experiments across twelve representative systems, multi-benchmark workload evaluations and cost-performance analyses reveal that memory architectures must align with task-specific bottlenecks and quantify each module’s contribution to overall performance. Key findings indicate no single architecture universally dominates; instead, localized maintenance strategies consistently outperform global reorganization, offering superior cost-effectiveness and guiding the design of agent-native memory systems.

agent memorydata managementdynamic knowledge update

We identify and prove a fundamental trade-off governing long-sequence models: no model can simultaneously achieve (i) per-step computation independent of sequence length (Efficiency), (ii) state size independent of sequence length (Compactness), and (iii) the ability to recall a number of historical facts proportional to sequence length (Recall). We formalize this trade-off within an Online Sequence Processor abstraction that unifies Transformers, state space models, linear recurrent networks, and their hybrids. Using the Data Processing Inequality and Fano's Inequality, we prove that any model satisfying Efficiency and Compactness can recall at most O(poly(d)/log V) key-value pairs from a sequence of arbitrary length, where d is the model dimension and V is the vocabulary size. We classify 52 architectures published before March 2026 into the triangle, showing that each achieves at most two of the three properties and that hybrid architectures trace continuous trajectories in the interior. Experiments on synthetic associative recall tasks with five representative architectures validate the theoretical bound: empirical recall capacity lies strictly below the information-theoretic limit, and no architecture escapes the triangle.

compactnessefficiencyimpossibility triangle

Hot Scholars

XL

Xin Lin

University of California San Diego
Computer VisionDeep Learning
XS

Xing Sun

Tencent Youtu Lab
LLMMLLMAgent
HJ

Heng Ji

Professor of Computer Science, AICE Director, ASKS Director, UIUC, Amazon Scholar
Natural Language ProcessingLarge Language Models