Score
Designs, implements, and analyzes machine learning models and agents that incorporate explicit memory structures and mechanisms—such as episodic stores, vector-indexed memories, translation-memory systems, and generative-recall modules—including the read/write/retrieval architectures and interfaces that connect memory to models. Develops training, regularization, and algorithmic approaches to mitigate catastrophic forgetting, tune precision/recall of retrieval, model recall behavior, and measure or analyze model memorization and memory movement.
This paper addresses the lack of formal modeling of memory mechanisms in LLM-based agents. Methodologically, it introduces the first unified, dynamic memory analysis framework, decomposing memory into three orthogonal representations—parametric, structured, and unstructured—and six atomic operations: consolidation, update, indexing, forgetting, retrieval, and compression. Through a systematic literature review and operation–theme mapping, the framework unifies research strands including long-term memory, long-context modeling, parameter editing, and retrieval-augmented generation (RAG). Its contributions are threefold: (1) construction of a comprehensive knowledge graph spanning all memory dimensions, integrating over 100 methods, benchmarks, and tools; (2) precise functional characterization and coordination pathways for each atomic operation; and (3) the first formal memory modeling foundation for LLM agents—enabling interpretable, scalable memory system design with both theoretical grounding and practical guidance.
This work addresses the challenges faced by foundational agents in long-horizon, dynamic, and user-dependent environments—particularly context explosion and sustained information management—which necessitate efficient memory mechanisms to enhance practical utility. The paper proposes the first unified three-dimensional framework that integrates internal and external memory, five cognitive mechanisms, and a dual-agent/user-centric perspective, offering a systematic structuring of research on agent memory. Drawing on a comprehensive review of hundreds of studies published before 2026, the framework synthesizes insights from memory modeling, cognitive science, and agent architecture to clarify memory operation strategies, evaluation benchmarks, and learning methodologies. It not only delineates structured pathways in existing research but also identifies key open problems, thereby providing a theoretical foundation and directional guidance for the future design of intelligent agent memory systems.
Current research on memory mechanisms in large language model (LLM) agents is fragmented across operating systems engineering and cognitive science, lacking a unified evolutionary perspective. This work proposes a three-stage memory evolution framework—storage, reflection, and experience—that systematically integrates recent advances in the field and formally defines the core drivers and key capabilities of each stage, such as active exploration and cross-trajectory abstraction. By synthesizing theoretical insights from cognitive science and systems engineering through comprehensive review and framework-based modeling, this study establishes a unified evolutionary theory of memory for LLM agents. The resulting framework offers clear design principles and a developmental roadmap for next-generation agents, advancing memory systems from passive recording toward active experience generation.
Current AI memory systems conflate retrieval with memory, relying solely on rapid exemplar storage while lacking the brain-inspired slow-weight integration mechanism, thereby limiting long-term learning, imposing theoretical ceilings on compositional generalization, and rendering them vulnerable to memory poisoning attacks. This work, grounded in the complementary learning systems theory from cognitive neuroscience, introduces a novel framework that formally distinguishes between “scratchpad” and “true memory,” rigorously analyzes the limitations of existing architectures, and incorporates a hippocampal–neocortical dual-system model. The study demonstrates that purely retrieval-based agents face an insurmountable generalization ceiling on novel compositional tasks and, building on this insight, proposes a coexisting memory architecture along with design principles to lay the theoretical foundation for next-generation AI memory systems capable of abstract rule-based generalization.
Current AI agent memory research is highly fragmented: conceptual definitions are ambiguous, taxonomic criteria inconsistent, and evaluation frameworks absent; the traditional dichotomy of long-term versus short-term memory fails to capture its heterogeneity. This paper addresses these gaps via a systematic literature review and conceptual modeling, proposing the first three-dimensional analytical framework for agent memory—unified along dimensions of *form* (token-level, parametric, implicit), *function* (factual, experiential, working), and *dynamics* (formation, evolution, retrieval)—while rigorously distinguishing agent memory from LLM internal memory, RAG, and context engineering. Contributions include: (1) a consensus-driven terminology system and taxonomy; (2) the first evaluation paradigm covering all three dimensions; (3) formal recognition of memory as a core agent capability; and (4) identification of key research directions—memory automation, multimodal coordination, multi-agent memory sharing, and trustworthiness—providing foundational support for architectural design and empirical investigation.
Existing benchmarks for long-term memory evaluation primarily focus on factual recall and simple retrieval, failing to adequately assess large language models’ (LLMs’) capacity to organize and leverage complex memory structures. To address this gap, this work introduces StructMemEval, a novel benchmark that systematically evaluates LLM agents’ ability to construct and utilize structured long-term memory in complex reasoning scenarios. StructMemEval emphasizes structural organization through tasks such as transaction ledgers, to-do lists, and tree-structured data. Experimental results demonstrate that mainstream LLMs struggle to autonomously organize memories into coherent structures without explicit guidance, whereas memory-augmented agents equipped with structural prompts achieve significantly higher task success rates. This benchmark thus fills a critical void in the evaluation of sophisticated memory architectures for LLM-based agents.
This work addresses the challenge that large language model (LLM) agents face in long-term interactions due to limited context windows, which hinder effective retention and utilization of past experiences. The authors propose a memory loop framework centered on “write–manage–read” and introduce the first three-dimensional taxonomy encompassing temporal scope, representational medium, and control strategy, systematically unifying five core families of memory mechanisms. By integrating techniques such as context compression, retrieval-augmented storage, reflective self-improvement, hierarchical virtual context, and policy-driven memory management, the study advances evaluation paradigms from static recall toward dynamic assessment of memory-decision integration across multi-session tasks. Analysis of four emerging benchmarks reveals current limitations in long-term memory consolidation and causal reasoning, while case studies in personal assistants, coding agents, and open-world gaming underscore the critical role of structured memory mechanisms.
This work addresses the construction of human-like memory mechanisms in large language models and multimodal foundation models to support continual learning, personalized reasoning, and cross-modal consistency. It proposes the first unified taxonomic framework that systematically integrates three major memory paradigms: implicit memory (e.g., parameterized memory), explicit memory (e.g., external retrieval and graph-structured knowledge bases), and agent memory. The framework is further extended to multimodal settings, elucidating the critical role of memory in cross-modal alignment and agent collaboration. Through a comprehensive literature review and taxonomic analysis, the study surveys existing technical approaches, evaluation benchmarks, and core challenges, thereby establishing a theoretical foundation and offering clear directions for future research on human-like memory systems in artificial intelligence.
Current research on memory mechanisms in large language models remains highly fragmented and lacks a unified theoretical framework. This work proposes an architecture-centered taxonomy that systematically models memory along three orthogonal dimensions: representation, update dynamics, and persistence, formally characterizing core processes such as writing, routing, state transition, and integration. By constructing the first three-dimensional unified framework that integrates implicit/explicit, offline/online, and short-term/long-term memory, it clarifies the boundary between computationally coupled memory and independently addressable memory. Through a systematic literature review, architectural analysis, and multidimensional evaluation, the study synthesizes techniques including attention mechanisms, recurrent states, parameter-efficient fine-tuning, and scalable retrieval-augmented storage, thereby establishing a coherent paradigm for memory modeling and providing a theoretical foundation and design principles for future scalable and adaptive large language models.
This work addresses the lack of effective long-term memory management in large language model agents during extended interactions. Inspired by human cognition, the authors propose a novel memory architecture that systematically integrates six key mechanisms: sleep-based consolidation, interference-driven forgetting, memory trace maturation, retrieval-induced reconsolidation, entity-centric knowledge graphs, and multi-cue hybrid retrieval. To prevent data leakage, they introduce an unsupervised synthetic calibration method for threshold setting. The framework further incorporates memory deduplication and compression, context budget control, and a streaming multi-level evaluation paradigm. Evaluated on the VSCode dataset, the approach achieves 97.2% memory retention accuracy while reducing storage overhead by 58%. On the LongMemEval benchmark, it matches baseline retrieval accuracy using only 200K tokens of context and improves S-tier preference recall by 13.3 percentage points.
This work addresses the efficiency challenges of cross-session memory management in large language model agents during long-horizon tasks, a domain where existing systems lack systematic understanding of memory mechanisms. The study introduces, for the first time, a four-dimensional taxonomy for agent memory and a phase-aware performance profiling framework. Leveraging two benchmark suites, the authors empirically evaluate ten representative systems, uncovering cost distribution patterns across memory construction, retrieval, and generation phases. The analysis quantifies trade-offs in write and read path overheads across different architectures and distills ten system design principles encompassing construction scheduling, capability baselines, query amortization, freshness–latency balance, and cluster management.
This study addresses the lack of fine-grained, system-level evaluation in existing memory systems for large language model agents, which hinders robust assessment under cost constraints, architectural trade-offs, and dynamic knowledge updates. The work proposes a unified analytical framework by decomposing agent memory into four independently evaluable core modules: representation storage, extraction, retrieval routing, and maintenance. Through end-to-end and ablation experiments across twelve representative systems, multi-benchmark workload evaluations and cost-performance analyses reveal that memory architectures must align with task-specific bottlenecks and quantify each module’s contribution to overall performance. Key findings indicate no single architecture universally dominates; instead, localized maintenance strategies consistently outperform global reorganization, offering superior cost-effectiveness and guiding the design of agent-native memory systems.
Current large language models rely on implicit memory, which hinders their capacity to support the advanced cognitive functions essential for artificial general intelligence (AGI), such as long-term planning, metacognition, and symbolic reasoning. This work systematically argues, for the first time, that explicit memory constitutes a critical pathway toward achieving AGI. Inspired by the neural mechanisms of the hippocampus, the study proposes a novel computational paradigm of artificial explicit memory that integrates insights from neuroscience with large language model architectures. The proposed framework not only establishes a theoretical foundation for designing explicit memory modules in AGI systems but also opens new interdisciplinary avenues at the intersection of cognitive science, neuroscience, and artificial intelligence.