memory mechanism design

Designs and evaluates mechanisms that store, update, retrieve, and manage information in computational systems, including temporal dynamics such as short‑term and long‑term storage, forgetting, consolidation, and time-dependent read/write behaviors. Work includes specifying architectures, algorithms, and interfaces for memory modules and analyzing their capacity, latency, consistency, and temporal access/retention properties.

memorymechanismdesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Current research on memory mechanisms in large language model (LLM) agents is fragmented across operating systems engineering and cognitive science, lacking a unified evolutionary perspective. This work proposes a three-stage memory evolution framework—storage, reflection, and experience—that systematically integrates recent advances in the field and formally defines the core drivers and key capabilities of each stage, such as active exploration and cross-trajectory abstraction. By synthesizing theoretical insights from cognitive science and systems engineering through comprehensive review and framework-based modeling, this study establishes a unified evolutionary theory of memory for LLM agents. The resulting framework offers clear design principles and a developmental roadmap for next-generation agents, advancing memory systems from passive recording toward active experience generation.

continual learningevolutionary frameworkLLM agent

Rethinking Memory in AI: Taxonomy, Operations, Topics, and Future Directions

May 01, 2025
YD
Yiming Du
🏛️ The Chinese University of Hong Kong | The University of Edinburgh | HKUST | Huawei UK R&D Ltd.

This paper addresses the lack of formal modeling of memory mechanisms in LLM-based agents. Methodologically, it introduces the first unified, dynamic memory analysis framework, decomposing memory into three orthogonal representations—parametric, structured, and unstructured—and six atomic operations: consolidation, update, indexing, forgetting, retrieval, and compression. Through a systematic literature review and operation–theme mapping, the framework unifies research strands including long-term memory, long-context modeling, parameter editing, and retrieval-augmented generation (RAG). Its contributions are threefold: (1) construction of a comprehensive knowledge graph spanning all memory dimensions, integrating over 100 methods, benchmarks, and tools; (2) precise functional characterization and coordination pathways for each atomic operation; and (3) the first formal memory modeling foundation for LLM agents—enabling interpretable, scalable memory system design with both theoretical grounding and practical guidance.

Classify memory representations in AI systemsIntroduce six fundamental memory operationsMap memory operations to key research topics

Must-Read Papers

Most classic and influential ideas
View more

Temporal Model On Quantum Logic

Feb 09, 2025
FD
Francesco D'Agostino

Existing models of temporal memory dynamics lack a unified formal framework integrating logical specification, cognitive principles, and probabilistic updating. Method: We propose “Temporal Memory Dynamics” (TMD): a formal theory unifying linear and branching temporal logic (LTL/CTL) for propositional evolution; incorporating exponential decay and Bayesian reactivation to model forgetting and recall; introducing a feedback-driven recursive memory chain; representing hierarchical memory and interference via directed acyclic graphs (DAGs); and defining an entropy-based recall efficiency metric. Contribution/Results: TMD is the first framework to systematically bridge temporal logic, cognitive forgetting theory, and Bayesian dynamic updating under formal verifiability and computational tractability. It significantly enhances theoretical consistency and explanatory power regarding memory decay, context-dependent retrieval, and interference suppression. By unifying symbolic, probabilistic, and structural modeling paradigms, TMD provides a scalable foundational model for cognitive modeling and neuromorphic computing.

Incorporate forgetting and reactivation mechanismsModel temporal memory dynamicsRepresent hierarchical memory organization

The First Principle of Big Memory Systems

Sep 30, 2023
YH
Yu Hua
🏛️ Huazhong University of Science and Technology

This paper addresses the underappreciated role of persistence in large-memory systems, asserting that “persistence must be treated as a first principle of system design.” It systematically analyzes structural challenges arising from vertical (capacity/latency) and horizontal (network flattening) scaling of the memory hierarchy. The authors propose a full-stack co-designed persistence paradigm: (1) elevating persistence to the foundational axiom of system architecture, and (2) introducing a dual-mechanism persistence model—predictable speculative persistence for performance and strongly consistent deterministic persistence for correctness. By enabling mobile persistence and jointly optimizing memory, interconnect, and storage layers, the approach achieves high throughput, low latency, and cost efficiency simultaneously. Extensive evaluations across diverse workloads demonstrate significant improvements in both system performance and reliability.

Achieving cost efficiency and high performance persistenceAnalyzing vertical and horizontal memory hierarchy extensionsExploring flattening networks in traditional storage hierarchies

Deep learning for time series has progressed through successive architectural paradigms, from recurrent networks and transformers to structured state-space models, retrieval-augmented predictors, foundation models, and tool-using agents. These developments are typically studied in isolation, organized by architecture or modeling era. We argue that they can instead be viewed through a common question of \emph{how does a time-series model retain and access information beyond its immediate input?} This question is motivated by a fundamental limitation of conventional time-series modeling: information relevant to a prediction may lie far beyond a feasible input window, while compressing history into a fixed-size state can discard information that may become useful later. We formulate this challenge as a \emph{memory} problem and organize existing time-series methods along a spectrum from internal memory, encoded in parameters and fixed-size states, to external memory that is addressable, retrievable, and increasingly maintained by agents. We then develop a unified taxonomy of memory mechanisms and review three classes of external memory, including explicit modules, retrieval augmentation, and agentic stores, under a common framework for what is retained, how it is written and accessed, and how it persists. A cross-cutting analysis maps these mechanisms to time-series tasks and identifies gaps in both methods and evaluation. We conclude by outlining open problems in building memory systems that can selectively retain, retrieve, revise, and forget information as temporal environments evolve. The result is a framework for studying memory as a first-class dimension of time series modeling, independent of the underlying backbone.

external memoryinformation retentionmemory problem

Performance Models for a Two-tiered Storage System

Mar 12, 2025
AS
Aparna Sasidharan
🏛️ IIT | Sandia National Lab | Oak Ridge National Lab

To address inefficient data migration and inaccurate performance prediction in heterogeneous storage systems (NVMe cache + HDD backend), this paper designs and implements a distributed two-tier storage system. We propose an online reinforcement learning–based dynamic data tiering scheduling algorithm and develop an end-to-end performance model integrating queuing network theory with fine-grained device behavior modeling. Our key contribution is the first scalable, fine-grained device behavior modeling method tailored for heterogeneous storage—enabling adaptive tiering management and precise performance prediction under high-concurrency I/O workloads in multi-core clusters. Experimental evaluation on multi-node clusters demonstrates an average model prediction error of less than 8%, a 27% improvement in I/O throughput, and a 34% reduction in average access latency. The framework provides a reusable modeling and optimization foundation for two-tier storage systems.

Design and analyze a two-tiered storage systemDevelop online learning for data tier managementEvaluate performance using queuing and behavioral models

This work addresses the efficiency challenges of cross-session memory management in large language model agents during long-horizon tasks, a domain where existing systems lack systematic understanding of memory mechanisms. The study introduces, for the first time, a four-dimensional taxonomy for agent memory and a phase-aware performance profiling framework. Leveraging two benchmark suites, the authors empirically evaluate ten representative systems, uncovering cost distribution patterns across memory construction, retrieval, and generation phases. The analysis quantifies trade-offs in write and read path overheads across different architectures and distills ten system design principles encompassing construction scheduling, capability baselines, query amortization, freshness–latency balance, and cluster management.

Agent MemoryLong-Horizon TasksMemory Systems

Latest Papers

What's happening recently
View more

Current research on memory mechanisms in large language models remains highly fragmented and lacks a unified theoretical framework. This work proposes an architecture-centered taxonomy that systematically models memory along three orthogonal dimensions: representation, update dynamics, and persistence, formally characterizing core processes such as writing, routing, state transition, and integration. By constructing the first three-dimensional unified framework that integrates implicit/explicit, offline/online, and short-term/long-term memory, it clarifies the boundary between computationally coupled memory and independently addressable memory. Through a systematic literature review, architectural analysis, and multidimensional evaluation, the study synthesizes techniques including attention mechanisms, recurrent states, parameter-efficient fine-tuning, and scalable retrieval-augmented storage, thereby establishing a coherent paradigm for memory modeling and providing a theoretical foundation and design principles for future scalable and adaptive large language models.

architectural paradigmsfragmentationlarge language models

This study addresses critical security risks—such as memory tampering, unauthorized access, and cross-session poisoning—that threaten the long-term memory of large language model (LLM) agents, an area insufficiently explored in existing research with regard to governability. Integrating insights from cognitive neuroscience and philosophy of memory, the work proposes a comprehensive memory lifecycle framework encompassing writing, storage, retrieval, execution, sharing, and forgetting. Aligning this framework with four core security objectives—integrity, confidentiality, availability, and governability—it conceptualizes memory as an independent security dimension and introduces the notion of “memory sovereignty,” identifying nine governance primitives. Through interdisciplinary analysis, the paper systematically categorizes threats including memory poisoning, extraction attacks, and control-flow hijacking, exposes gaps in current architectures concerning governance coverage and LLMs’ reflexive security capabilities, and highlights the untapped potential of leveraging LLMs themselves to enforce memory security.

LLM agentslong-term memorymemory security

Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.

distributed systemsparallel systemsruntime control

Current AI agents treat long-term memory as static records, leading to uncontrolled growth, lack of semantic revision, capacity-driven forgetting, and read-only retrieval—limitations that undermine memory efficacy. This work proposes the Governed Evolving Memory (GEM) framework, which reconceptualizes long-term memory as a trajectory of states rather than isolated snapshots. GEM introduces state-level correctness conditions and defines four state-aware operations: ingestion, revision, forgetting, and retrieval. The authors implement MemState, a prototype system based on property graphs, to demonstrate GEM’s feasibility. Their evaluation reveals a fundamental gap between conventional database engines and memory-native architectures, establishing a memory-centric paradigm for data management and charting directions for future research in intelligent agent memory systems.

AI agent memorydata managementlong-term memory

This work challenges the prevailing assumption that large language model (LLM) agents inherently function without structured long-term memory, presenting the first systematic investigation into file systems as a memory substrate. The authors introduce a unified memory architecture comprising three collaborating agent types—manager, searcher, and executor—that jointly maintain a directory-tree-based Markdown storage system augmented with sandboxed shell tools, dedicated memory interfaces, and chunked retrieval strategies. Through comprehensive evaluation, they assess how organizational schemes, tooling, and model capabilities jointly influence memory coherence and task performance. Results demonstrate that well-structured memory organization can reduce retrieval costs by approximately 50%, yet most agents struggle to sustain such structure over time. Notably, changes to the toolset exert an impact on memory morphology comparable to switching the underlying LLM, underscoring that memory architecture constitutes a design space rather than a fixed default.

filesystem-based memoryLLM agentslong-term memory

Hot Scholars

TS

Tianyu Shi

University of Toronto
Reinforcement learningIntelligent Transportation SystemLarge Language ModelsAI
XH

Xianglong He

Tsinghua University
3D GenerationVideo GenerationMLLM
ZX

Zhiming Xu

University of Virginia
llm inferencemachine learning system
JZ

Jiahuan Zhou

Peking University
Computer VisionMachine LearningDeep Learning
JT

Jin Tang

Anhui University
Computer visionintelligent video analysis