๐ค AI Summary
This work addresses the challenge that large language model (LLM) agents face in long-term interactions due to limited context windows, which hinder effective retention and utilization of past experiences. The authors propose a memory loop framework centered on โwriteโmanageโreadโ and introduce the first three-dimensional taxonomy encompassing temporal scope, representational medium, and control strategy, systematically unifying five core families of memory mechanisms. By integrating techniques such as context compression, retrieval-augmented storage, reflective self-improvement, hierarchical virtual context, and policy-driven memory management, the study advances evaluation paradigms from static recall toward dynamic assessment of memory-decision integration across multi-session tasks. Analysis of four emerging benchmarks reveals current limitations in long-term memory consolidation and causal reasoning, while case studies in personal assistants, coding agents, and open-world gaming underscore the critical role of structured memory mechanisms.
๐ Abstract
Large language model (LLM) agents increasingly operate in settings where a single context window is far too small to capture what has happened, what was learned, and what should not be repeated. Memory -- the ability to persist, organize, and selectively recall information across interactions -- is what turns a stateless text generator into a genuinely adaptive agent. This survey offers a structured account of how memory is designed, implemented, and evaluated in modern LLM-based agents, covering work from 2022 through early 2026. We formalize agent memory as a \emph{write--manage--read} loop tightly coupled with perception and action, then introduce a three-dimensional taxonomy spanning temporal scope, representational substrate, and control policy. Five mechanism families are examined in depth: context-resident compression, retrieval-augmented stores, reflective self-improvement, hierarchical virtual context, and policy-learned management. On the evaluation side, we trace the shift from static recall benchmarks to multi-session agentic tests that interleave memory with decision-making, analyzing four recent benchmarks that expose stubborn gaps in current systems. We also survey applications where memory is the differentiating factor -- personal assistants, coding agents, open-world games, scientific reasoning, and multi-agent teamwork -- and address the engineering realities of write-path filtering, contradiction handling, latency budgets, and privacy governance. The paper closes with open challenges: continual consolidation, causally grounded retrieval, trustworthy reflection, learned forgetting, and multimodal embodied memory.