agent memory systems

Designs, builds, and evaluates memory architectures for agents—including LLM-based agents—that store persistent episodic and semantic representations (external memory modules, summaries, traces, and indices) to enable runtime retrieval and reuse of past interactions. This work covers implementing retrieval mechanisms (e.g., RAG-style and query interfaces), synchronization and persistence across agent instances, episodic structuring and replay policies (including cross‑episode or cross‑season reuse), and memory update/management strategies that bias reasoning or reduce compute by reusing past solutions without model retraining.

agentmemorysystems

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.57
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$211K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Current research on memory mechanisms in large language model (LLM) agents is fragmented across operating systems engineering and cognitive science, lacking a unified evolutionary perspective. This work proposes a three-stage memory evolution framework—storage, reflection, and experience—that systematically integrates recent advances in the field and formally defines the core drivers and key capabilities of each stage, such as active exploration and cross-trajectory abstraction. By synthesizing theoretical insights from cognitive science and systems engineering through comprehensive review and framework-based modeling, this study establishes a unified evolutionary theory of memory for LLM agents. The resulting framework offers clear design principles and a developmental roadmap for next-generation agents, advancing memory systems from passive recording toward active experience generation.

continual learningevolutionary frameworkLLM agent

Current agent memory systems lack reliable evaluation benchmarks and systematic architectural analysis, leading to distorted performance assessments and ad hoc design choices. This work addresses this gap by proposing, for the first time, a taxonomy grounded in four distinct memory architectures, examined through empirical studies across multiple large language model backbones. The study systematically evaluates these architectures in terms of semantic utility, benchmark saturation, model dependency, and memory overhead, uncovering fundamental limitations that explain why real-world performance consistently falls short of theoretical expectations. These findings provide critical empirical evidence and actionable insights for designing scalable, evaluable memory systems in artificial agents.

agentic memorybenchmark limitationsevaluation metrics

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge that large language model (LLM) agents face in long-term interactions due to limited context windows, which hinder effective retention and utilization of past experiences. The authors propose a memory loop framework centered on “write–manage–read” and introduce the first three-dimensional taxonomy encompassing temporal scope, representational medium, and control strategy, systematically unifying five core families of memory mechanisms. By integrating techniques such as context compression, retrieval-augmented storage, reflective self-improvement, hierarchical virtual context, and policy-driven memory management, the study advances evaluation paradigms from static recall toward dynamic assessment of memory-decision integration across multi-session tasks. Analysis of four emerging benchmarks reveals current limitations in long-term memory consolidation and causal reasoning, while case studies in personal assistants, coding agents, and open-world gaming underscore the critical role of structured memory mechanisms.

adaptive agentcontext windowinformation recall

Current memory mechanisms in large language model agents lack systematic comparison and a unified framework, hindering their ability to support long-horizon, complex tasks. This work proposes the first unified architecture encompassing mainstream memory approaches and conducts a comprehensive evaluation and ablation study under standardized benchmarks and consistent experimental conditions, revealing the strengths, weaknesses, and suitable application scenarios of each method. Building on these insights, we design a modular memory mechanism that combines optimized strategies to significantly outperform existing state-of-the-art methods. The effectiveness of our approach is validated on two established benchmarks, offering a promising new direction for future research on agent memory systems.

agent memorycomparative analysislarge language models

Memory in the Age of AI Agents

Dec 15, 2025
YH
Yuyang Hu
🏛️ Renmin University of China | Fudan University | Peking University | National University of Singapore

Current AI agent memory research is highly fragmented: conceptual definitions are ambiguous, taxonomic criteria inconsistent, and evaluation frameworks absent; the traditional dichotomy of long-term versus short-term memory fails to capture its heterogeneity. This paper addresses these gaps via a systematic literature review and conceptual modeling, proposing the first three-dimensional analytical framework for agent memory—unified along dimensions of *form* (token-level, parametric, implicit), *function* (factual, experiential, working), and *dynamics* (formation, evolution, retrieval)—while rigorously distinguishing agent memory from LLM internal memory, RAG, and context engineering. Contributions include: (1) a consensus-driven terminology system and taxonomy; (2) the first evaluation paradigm covering all three dimensions; (3) formal recognition of memory as a core agent capability; and (4) identification of key research directions—memory automation, multimodal coordination, multi-agent memory sharing, and trustworthiness—providing foundational support for architectural design and empirical investigation.

Clarifying memory scope and distinguishing from related conceptsProviding taxonomy and benchmarks for future agent developmentSurveying fragmented research on AI agent memory systems

This work addresses the efficiency challenges of cross-session memory management in large language model agents during long-horizon tasks, a domain where existing systems lack systematic understanding of memory mechanisms. The study introduces, for the first time, a four-dimensional taxonomy for agent memory and a phase-aware performance profiling framework. Leveraging two benchmark suites, the authors empirically evaluate ten representative systems, uncovering cost distribution patterns across memory construction, retrieval, and generation phases. The analysis quantifies trade-offs in write and read path overheads across different architectures and distills ten system design principles encompassing construction scheduling, capability baselines, query amortization, freshness–latency balance, and cluster management.

Agent MemoryLong-Horizon TasksMemory Systems

This work addresses the construction of human-like memory mechanisms in large language models and multimodal foundation models to support continual learning, personalized reasoning, and cross-modal consistency. It proposes the first unified taxonomic framework that systematically integrates three major memory paradigms: implicit memory (e.g., parameterized memory), explicit memory (e.g., external retrieval and graph-structured knowledge bases), and agent memory. The framework is further extended to multimodal settings, elucidating the critical role of memory in cross-modal alignment and agent collaboration. Through a comprehensive literature review and taxonomic analysis, the study surveys existing technical approaches, evaluation benchmarks, and core challenges, thereby establishing a theoretical foundation and offering clear directions for future research on human-like memory systems in artificial intelligence.

autonomous agentscontinual learningLarge Language Models

Latest Papers

What's happening recently
View more

This work addresses the challenges large language models face in long-horizon tasks regarding knowledge retention, organization, and reuse. Existing approaches are often limited to static fact storage or replay of successful experiences, struggling to incorporate failure cases and lacking online extensibility. To overcome these limitations, the paper proposes a unified memory framework that, for the first time, integrates semantic, episodic, and procedural memory within a dual-layer short- and long-term storage architecture. A multi-agent design—comprising actor, memory, and critic components—enables automatic memory generation, reward-based annotation, and adaptive retrieval. A reward-driven long-term memory management strategy, including evaluation, consolidation, and pruning, facilitates continual learning and online evolution. Experiments demonstrate that the proposed method significantly improves success rates and robustness across diverse complex long-horizon tasks, outperforming current baselines.

LLM-based agentslong-horizon tasksmemory framework

Existing benchmarks for long-term memory evaluation primarily focus on factual recall and simple retrieval, failing to adequately assess large language models’ (LLMs’) capacity to organize and leverage complex memory structures. To address this gap, this work introduces StructMemEval, a novel benchmark that systematically evaluates LLM agents’ ability to construct and utilize structured long-term memory in complex reasoning scenarios. StructMemEval emphasizes structural organization through tasks such as transaction ledgers, to-do lists, and tree-structured data. Experimental results demonstrate that mainstream LLMs struggle to autonomously organize memories into coherent structures without explicit guidance, whereas memory-augmented agents equipped with structural prompts achieve significantly higher task success rates. This benchmark thus fills a critical void in the evaluation of sophisticated memory architectures for LLM-based agents.

benchmarkLLM agentslong-term memory

This study addresses the lack of fine-grained, system-level evaluation in existing memory systems for large language model agents, which hinders robust assessment under cost constraints, architectural trade-offs, and dynamic knowledge updates. The work proposes a unified analytical framework by decomposing agent memory into four independently evaluable core modules: representation storage, extraction, retrieval routing, and maintenance. Through end-to-end and ablation experiments across twelve representative systems, multi-benchmark workload evaluations and cost-performance analyses reveal that memory architectures must align with task-specific bottlenecks and quantify each module’s contribution to overall performance. Key findings indicate no single architecture universally dominates; instead, localized maintenance strategies consistently outperform global reorganization, offering superior cost-effectiveness and guiding the design of agent-native memory systems.

agent memorydata managementdynamic knowledge update

This work addresses the challenges of cross-session state persistence, user preference retention, and procedural knowledge accumulation in enterprise-grade, long-horizon AI agents by proposing a native memory architecture built upon Oracle Database. The architecture employs a hierarchical design that decouples active memory from passive storage, supports fine-grained scope control, and encompasses the full lifecycle of memory management. By integrating mechanisms for memory ingestion, retrieval, summarization, and revision, the system achieves 93.8% accuracy on the LongMemEval benchmark—significantly outperforming a flat-history baseline while reducing token consumption by approximately 10.7×. This approach substantially enhances memory efficiency and task performance without compromising low latency or high precision.

agent memoryenterprise memorylong-horizon AI agents

Existing memory systems based on semantic similarity struggle to maintain execution state consistency in long-horizon tasks, often leading to fragmented decision-making and error propagation. This work proposes MAGE, a novel framework that conceptualizes memory as an execution state manager. MAGE organizes interaction history into a hierarchical state tree, where the path from root to the current node defines the agent’s state, and integrates subgoal summaries, recent trajectories, and historical branch hints for decision-making. The framework introduces four coordinated operations—Grow, Compress, Maintain, and Revise—to dynamically manage the state tree, enabling effective error isolation and efficient contextual utilization. Evaluated on the MemoryArena benchmark, MAGE improves task success rates by 7.8–20.4 percentage points while reducing token consumption by 55.1%.

agent memoryerror isolationexecution-state dependencies

Hot Scholars

YW

Ying Wen

Associate Professor, Shanghai Jiao Tong University
Multi-Agent LearningReinforcement Learning
BT

Bo Tang

Hefei, Anhui Province, China
radar signal processingwaveform designMIMOISAC
MW

Muning Wen

Research Assistant Professor, Shanghai Jiao Tong University
(multi-agent) reinforcement learninglanguage agent/LLM-based agent
WZ

Weinan Zhang

Professor, Shanghai Jiao Tong University
Reinforcement LearningAgentsData Science
FX

Feiyu Xiong

MemTensor (Shanghai) Technology Co., Ltd.
Machine LearningNLPLLM