memory retrieval and querying

Design and build systems that index, query, and retrieve stored memory representations — including construction of index structures and lookup APIs, retrieval and ranking models, associative and graph-based activation mechanisms, and stable, efficient storage/lookup pathways. Implement and analyze retrieval workflows and agents that perform iterative or multi‑round retrieval and composition, agentic/automated retrieval, parallelized retrieval alongside memory construction, and verification/correction of retrieved items to supply compact, contextually relevant memory to downstream processes.

memoryretrievalandquerying

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$218K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Rethinking Memory in AI: Taxonomy, Operations, Topics, and Future Directions

May 01, 2025
YD
Yiming Du
🏛️ The Chinese University of Hong Kong | The University of Edinburgh | HKUST | Huawei UK R&D Ltd.

This paper addresses the lack of formal modeling of memory mechanisms in LLM-based agents. Methodologically, it introduces the first unified, dynamic memory analysis framework, decomposing memory into three orthogonal representations—parametric, structured, and unstructured—and six atomic operations: consolidation, update, indexing, forgetting, retrieval, and compression. Through a systematic literature review and operation–theme mapping, the framework unifies research strands including long-term memory, long-context modeling, parameter editing, and retrieval-augmented generation (RAG). Its contributions are threefold: (1) construction of a comprehensive knowledge graph spanning all memory dimensions, integrating over 100 methods, benchmarks, and tools; (2) precise functional characterization and coordination pathways for each atomic operation; and (3) the first formal memory modeling foundation for LLM agents—enabling interpretable, scalable memory system design with both theoretical grounding and practical guidance.

Classify memory representations in AI systemsIntroduce six fundamental memory operationsMap memory operations to key research topics

This work addresses the challenges faced by foundational agents in long-horizon, dynamic, and user-dependent environments—particularly context explosion and sustained information management—which necessitate efficient memory mechanisms to enhance practical utility. The paper proposes the first unified three-dimensional framework that integrates internal and external memory, five cognitive mechanisms, and a dual-agent/user-centric perspective, offering a systematic structuring of research on agent memory. Drawing on a comprehensive review of hundreds of studies published before 2026, the framework synthesizes insights from memory modeling, cognitive science, and agent architecture to clarify memory operation strategies, evaluation benchmarks, and learning methodologies. It not only delineates structured pathways in existing research but also identifies key open problems, thereby providing a theoretical foundation and directional guidance for the future design of intelligent agent memory systems.

context explosionfoundation agentslong-horizon environments

Must-Read Papers

Most classic and influential ideas
View more

Current AI memory systems conflate retrieval with memory, relying solely on rapid exemplar storage while lacking the brain-inspired slow-weight integration mechanism, thereby limiting long-term learning, imposing theoretical ceilings on compositional generalization, and rendering them vulnerable to memory poisoning attacks. This work, grounded in the complementary learning systems theory from cognitive neuroscience, introduces a novel framework that formally distinguishes between “scratchpad” and “true memory,” rigorously analyzes the limitations of existing architectures, and incorporates a hippocampal–neocortical dual-system model. The study demonstrates that purely retrieval-based agents face an insurmountable generalization ceiling on novel compositional tasks and, building on this insight, proposes a coexisting memory architecture along with design principles to lay the theoretical foundation for next-generation AI memory systems capable of abstract rule-based generalization.

agentic memorycomplementary learning systemsgeneralization

This work addresses the efficiency challenges of cross-session memory management in large language model agents during long-horizon tasks, a domain where existing systems lack systematic understanding of memory mechanisms. The study introduces, for the first time, a four-dimensional taxonomy for agent memory and a phase-aware performance profiling framework. Leveraging two benchmark suites, the authors empirically evaluate ten representative systems, uncovering cost distribution patterns across memory construction, retrieval, and generation phases. The analysis quantifies trade-offs in write and read path overheads across different architectures and distills ten system design principles encompassing construction scheduling, capability baselines, query amortization, freshness–latency balance, and cluster management.

Agent MemoryLong-Horizon TasksMemory Systems

This study addresses the relative importance of memory writing and retrieval in large language model agents and identifies the primary performance bottleneck. To this end, we propose the first diagnostic framework that disentangles bottlenecks across memory writing, retrieval, and utilization stages. We conduct a 3×3 controlled experiment on the LoCoMo benchmark, evaluating three writing strategies—naive chunking, Mem0-style fact extraction, and MemGPT-style summarization—paired with three retrieval methods: cosine similarity, BM25, and hybrid re-ranking. Results reveal that retrieval methods dominate performance variation, yielding up to a 20-percentage-point accuracy gap, whereas writing strategies exert limited influence. Notably, naive chunking—requiring no LLM calls—matches or surpasses more complex writing approaches, indicating that current systems are primarily constrained by retrieval efficacy rather than writing sophistication.

LLM agent memorymemory utilizationretrieval bottleneck

Current AI agents predominantly rely on vector-based retrieval for long-term memory, which suffers from opacity, limited auditability, and constrained expressive power. This work proposes the first fully auditable, model-agnostic, and storage-decoupled structured relational memory architecture, driven by the agent itself. The system enables transparent memory management through symbolic querying, automated concept induction, and a metadata knowledge graph—without requiring external ontologies. Implemented atop a relational database, it was evaluated on a single-user academic corpus over one year, processing 44 million tokens, 110,000 text segments, 160,000 documents, and approximately 5 million relations, thereby demonstrating strong scalability and long-term stability.

agentic memoryauditable AIlong-term memory

This work challenges the prevailing assumption that large language model (LLM) agents inherently function without structured long-term memory, presenting the first systematic investigation into file systems as a memory substrate. The authors introduce a unified memory architecture comprising three collaborating agent types—manager, searcher, and executor—that jointly maintain a directory-tree-based Markdown storage system augmented with sandboxed shell tools, dedicated memory interfaces, and chunked retrieval strategies. Through comprehensive evaluation, they assess how organizational schemes, tooling, and model capabilities jointly influence memory coherence and task performance. Results demonstrate that well-structured memory organization can reduce retrieval costs by approximately 50%, yet most agents struggle to sustain such structure over time. Notably, changes to the toolset exert an impact on memory morphology comparable to switching the underlying LLM, underscoring that memory architecture constitutes a design space rather than a fixed default.

filesystem-based memoryLLM agentslong-term memory

Latest Papers

What's happening recently
View more

This study addresses the lack of fine-grained, system-level evaluation in existing memory systems for large language model agents, which hinders robust assessment under cost constraints, architectural trade-offs, and dynamic knowledge updates. The work proposes a unified analytical framework by decomposing agent memory into four independently evaluable core modules: representation storage, extraction, retrieval routing, and maintenance. Through end-to-end and ablation experiments across twelve representative systems, multi-benchmark workload evaluations and cost-performance analyses reveal that memory architectures must align with task-specific bottlenecks and quantify each module’s contribution to overall performance. Key findings indicate no single architecture universally dominates; instead, localized maintenance strategies consistently outperform global reorganization, offering superior cost-effectiveness and guiding the design of agent-native memory systems.

agent memorydata managementdynamic knowledge update

Existing language agents lack a unified formal framework for systematically comparing hierarchical memory designs. This work proposes the first formal theory that models the entire hierarchical memory process—from raw data to multi-granularity representations and constrained context retrieval—through three operators: extraction (α), coarsening (C), and traversal (τ). The framework defines a spectrum of self-contained memory representations, reveals coupling constraints between coarsening and traversal strategies, and enables information extraction, multi-level clustered compression, and query- and budget-aware context traversal. Its generality and explanatory power are validated through application to 11 representative systems spanning document hierarchies, dialogue memory, and agent trajectories.

context-length limitationdesign comparisonformalism

Current memory systems for large language model agents are predominantly designed for single scenarios, limiting their generalization across diverse tasks. This work systematically evaluates eight memory mechanisms alongside a search-oriented memory framework across five heterogeneous environments to assess their universality. To address these limitations, the paper introduces AutoMEM, a novel mechanism enabling agents to autonomously manage memory storage and retrieval. Departing from conventional fixed-pipeline passive storage, AutoMEM employs flat textual memory representations accessed via tool calls, empowering agents with active control over when and how to store or retrieve memories. Experimental results demonstrate that AutoMEM significantly outperforms baseline approaches in cross-task overall performance, underscoring the critical role of autonomous memory management in enhancing agent generalization.

agentic memorycross-scenario generalityheterogeneous trajectories

Traditional language agents are constrained by high latency and limited reasoning capabilities due to infrequent access to external memory, hindering the sustained availability akin to human working memory. This work proposes embedding memory storage directly within the agent’s reasoning loop, constructing an end-to-end language agent system through in-process vector storage, lightweight local embedding models, and efficient read-write strategies. The study demonstrates for the first time that this architecture reduces retrieval latency to 80–165 microseconds (p50), significantly improving task recall accuracy from 0/5 to 3.6–4.8/5 under a fixed latency budget, while achieving zero data loss across 244 write operations. These results provide a scalable engineering realization supporting the extended mind hypothesis.

extended working memoryin-loop retrievallanguage agents

Hot Scholars

DL

Dongha Lee

Yonsei University
Data miningInformation retrievalNatural language processing
DK

Dmitry Krotov

MIT-IBM Watson AI Lab & IBM Research
Neural NetworksMachine LearningArtificial Intelligence
FX

Feiyu Xiong

MemTensor (Shanghai) Technology Co., Ltd.
Machine LearningNLPLLM
ZL

Zhiyu Li

Tianjin University
Robust controlattitude control
JM

Julian McAuley

Professor, UC San Diego
Recommender SystemsNatural Language ProcessingPersonalizationComputer Music