Score
Design and build systems that index, query, and retrieve stored memory representations — including construction of index structures and lookup APIs, retrieval and ranking models, associative and graph-based activation mechanisms, and stable, efficient storage/lookup pathways. Implement and analyze retrieval workflows and agents that perform iterative or multi‑round retrieval and composition, agentic/automated retrieval, parallelized retrieval alongside memory construction, and verification/correction of retrieved items to supply compact, contextually relevant memory to downstream processes.
Current agentic RAG systems lack a unified theoretical framework, resulting in architectural fragmentation, inconsistent evaluation protocols, and insufficiently characterized reliability risks. This work addresses these challenges by formally modeling agentic RAG as a finite-horizon partially observable Markov decision process. It introduces a modular architecture and a systematic taxonomy encompassing core components such as planning, retrieval coordination, memory paradigms, and tool invocation. By analyzing the limitations of static evaluation and identifying dynamic risks inherent in autonomous loops—particularly hallucination propagation and memory contamination—the study establishes a theoretical foundation for agentic RAG and proposes reliability-oriented evaluation criteria. Furthermore, it outlines key directions for future research, including adaptive retrieval strategies, cost-aware coordination mechanisms, and effective supervision frameworks.
This paper addresses the lack of formal modeling of memory mechanisms in LLM-based agents. Methodologically, it introduces the first unified, dynamic memory analysis framework, decomposing memory into three orthogonal representations—parametric, structured, and unstructured—and six atomic operations: consolidation, update, indexing, forgetting, retrieval, and compression. Through a systematic literature review and operation–theme mapping, the framework unifies research strands including long-term memory, long-context modeling, parameter editing, and retrieval-augmented generation (RAG). Its contributions are threefold: (1) construction of a comprehensive knowledge graph spanning all memory dimensions, integrating over 100 methods, benchmarks, and tools; (2) precise functional characterization and coordination pathways for each atomic operation; and (3) the first formal memory modeling foundation for LLM agents—enabling interpretable, scalable memory system design with both theoretical grounding and practical guidance.
This work addresses the challenges faced by foundational agents in long-horizon, dynamic, and user-dependent environments—particularly context explosion and sustained information management—which necessitate efficient memory mechanisms to enhance practical utility. The paper proposes the first unified three-dimensional framework that integrates internal and external memory, five cognitive mechanisms, and a dual-agent/user-centric perspective, offering a systematic structuring of research on agent memory. Drawing on a comprehensive review of hundreds of studies published before 2026, the framework synthesizes insights from memory modeling, cognitive science, and agent architecture to clarify memory operation strategies, evaluation benchmarks, and learning methodologies. It not only delineates structured pathways in existing research but also identifies key open problems, thereby providing a theoretical foundation and directional guidance for the future design of intelligent agent memory systems.
Current AI memory systems conflate retrieval with memory, relying solely on rapid exemplar storage while lacking the brain-inspired slow-weight integration mechanism, thereby limiting long-term learning, imposing theoretical ceilings on compositional generalization, and rendering them vulnerable to memory poisoning attacks. This work, grounded in the complementary learning systems theory from cognitive neuroscience, introduces a novel framework that formally distinguishes between “scratchpad” and “true memory,” rigorously analyzes the limitations of existing architectures, and incorporates a hippocampal–neocortical dual-system model. The study demonstrates that purely retrieval-based agents face an insurmountable generalization ceiling on novel compositional tasks and, building on this insight, proposes a coexisting memory architecture along with design principles to lay the theoretical foundation for next-generation AI memory systems capable of abstract rule-based generalization.
This work addresses the efficiency challenges of cross-session memory management in large language model agents during long-horizon tasks, a domain where existing systems lack systematic understanding of memory mechanisms. The study introduces, for the first time, a four-dimensional taxonomy for agent memory and a phase-aware performance profiling framework. Leveraging two benchmark suites, the authors empirically evaluate ten representative systems, uncovering cost distribution patterns across memory construction, retrieval, and generation phases. The analysis quantifies trade-offs in write and read path overheads across different architectures and distills ten system design principles encompassing construction scheduling, capability baselines, query amortization, freshness–latency balance, and cluster management.
This study addresses the relative importance of memory writing and retrieval in large language model agents and identifies the primary performance bottleneck. To this end, we propose the first diagnostic framework that disentangles bottlenecks across memory writing, retrieval, and utilization stages. We conduct a 3×3 controlled experiment on the LoCoMo benchmark, evaluating three writing strategies—naive chunking, Mem0-style fact extraction, and MemGPT-style summarization—paired with three retrieval methods: cosine similarity, BM25, and hybrid re-ranking. Results reveal that retrieval methods dominate performance variation, yielding up to a 20-percentage-point accuracy gap, whereas writing strategies exert limited influence. Notably, naive chunking—requiring no LLM calls—matches or surpasses more complex writing approaches, indicating that current systems are primarily constrained by retrieval efficacy rather than writing sophistication.
Current AI agents predominantly rely on vector-based retrieval for long-term memory, which suffers from opacity, limited auditability, and constrained expressive power. This work proposes the first fully auditable, model-agnostic, and storage-decoupled structured relational memory architecture, driven by the agent itself. The system enables transparent memory management through symbolic querying, automated concept induction, and a metadata knowledge graph—without requiring external ontologies. Implemented atop a relational database, it was evaluated on a single-user academic corpus over one year, processing 44 million tokens, 110,000 text segments, 160,000 documents, and approximately 5 million relations, thereby demonstrating strong scalability and long-term stability.
This work challenges the prevailing assumption that large language model (LLM) agents inherently function without structured long-term memory, presenting the first systematic investigation into file systems as a memory substrate. The authors introduce a unified memory architecture comprising three collaborating agent types—manager, searcher, and executor—that jointly maintain a directory-tree-based Markdown storage system augmented with sandboxed shell tools, dedicated memory interfaces, and chunked retrieval strategies. Through comprehensive evaluation, they assess how organizational schemes, tooling, and model capabilities jointly influence memory coherence and task performance. Results demonstrate that well-structured memory organization can reduce retrieval costs by approximately 50%, yet most agents struggle to sustain such structure over time. Notably, changes to the toolset exert an impact on memory morphology comparable to switching the underlying LLM, underscoring that memory architecture constitutes a design space rather than a fixed default.
This study addresses the lack of fine-grained, system-level evaluation in existing memory systems for large language model agents, which hinders robust assessment under cost constraints, architectural trade-offs, and dynamic knowledge updates. The work proposes a unified analytical framework by decomposing agent memory into four independently evaluable core modules: representation storage, extraction, retrieval routing, and maintenance. Through end-to-end and ablation experiments across twelve representative systems, multi-benchmark workload evaluations and cost-performance analyses reveal that memory architectures must align with task-specific bottlenecks and quantify each module’s contribution to overall performance. Key findings indicate no single architecture universally dominates; instead, localized maintenance strategies consistently outperform global reorganization, offering superior cost-effectiveness and guiding the design of agent-native memory systems.
Existing language agents lack a unified formal framework for systematically comparing hierarchical memory designs. This work proposes the first formal theory that models the entire hierarchical memory process—from raw data to multi-granularity representations and constrained context retrieval—through three operators: extraction (α), coarsening (C), and traversal (τ). The framework defines a spectrum of self-contained memory representations, reveals coupling constraints between coarsening and traversal strategies, and enables information extraction, multi-level clustered compression, and query- and budget-aware context traversal. Its generality and explanatory power are validated through application to 11 representative systems spanning document hierarchies, dialogue memory, and agent trajectories.
Current memory systems for large language model agents are predominantly designed for single scenarios, limiting their generalization across diverse tasks. This work systematically evaluates eight memory mechanisms alongside a search-oriented memory framework across five heterogeneous environments to assess their universality. To address these limitations, the paper introduces AutoMEM, a novel mechanism enabling agents to autonomously manage memory storage and retrieval. Departing from conventional fixed-pipeline passive storage, AutoMEM employs flat textual memory representations accessed via tool calls, empowering agents with active control over when and how to store or retrieve memories. Experimental results demonstrate that AutoMEM significantly outperforms baseline approaches in cross-task overall performance, underscoring the critical role of autonomous memory management in enhancing agent generalization.
Traditional language agents are constrained by high latency and limited reasoning capabilities due to infrequent access to external memory, hindering the sustained availability akin to human working memory. This work proposes embedding memory storage directly within the agent’s reasoning loop, constructing an end-to-end language agent system through in-process vector storage, lightweight local embedding models, and efficient read-write strategies. The study demonstrates for the first time that this architecture reduces retrieval latency to 80–165 microseconds (p50), significantly improving task recall accuracy from 0/5 to 3.6–4.8/5 under a fixed latency budget, while achieving zero data loss across 244 write operations. These results provide a scalable engineering realization supporting the extended mind hypothesis.