Score
Designs and evaluates protocols, metrics, and reconstruction methods that measure an agent’s stored scene or episodic memory independently of its action sequence — e.g., decoupling memory fidelity from actions, implementing multi-level or long-term (LTP) memory evaluations, and scoring memory over extended interactions. Builds transition-localized 3D reconstruction and subject-tracking pipelines and integrates vision–language reasoning to localize important transitions and assess memory quality at multiple granularities.
This work addresses the challenges faced by foundational agents in long-horizon, dynamic, and user-dependent environments—particularly context explosion and sustained information management—which necessitate efficient memory mechanisms to enhance practical utility. The paper proposes the first unified three-dimensional framework that integrates internal and external memory, five cognitive mechanisms, and a dual-agent/user-centric perspective, offering a systematic structuring of research on agent memory. Drawing on a comprehensive review of hundreds of studies published before 2026, the framework synthesizes insights from memory modeling, cognitive science, and agent architecture to clarify memory operation strategies, evaluation benchmarks, and learning methodologies. It not only delineates structured pathways in existing research but also identifies key open problems, thereby providing a theoretical foundation and directional guidance for the future design of intelligent agent memory systems.
Current research on memory mechanisms in large language model (LLM) agents is fragmented across operating systems engineering and cognitive science, lacking a unified evolutionary perspective. This work proposes a three-stage memory evolution framework—storage, reflection, and experience—that systematically integrates recent advances in the field and formally defines the core drivers and key capabilities of each stage, such as active exploration and cross-trajectory abstraction. By synthesizing theoretical insights from cognitive science and systems engineering through comprehensive review and framework-based modeling, this study establishes a unified evolutionary theory of memory for LLM agents. The resulting framework offers clear design principles and a developmental roadmap for next-generation agents, advancing memory systems from passive recording toward active experience generation.
Existing action-conditioned world models often suffer from silent scene changes after the camera departs due to memory failure, and lack a standardized benchmark for fairly comparing memory mechanisms. This work introduces a controlled experimental framework that systematically disentangles four key dimensions of memory—capacity, compression, retrieval, and recurrence—while fixing the video diffusion backbone, optimizer, action representation, and evaluation protocol. A three-branch evaluation scheme reveals that replay fidelity alone is insufficient to assess true memory capability. By unifying the action-to-video interface to compare mechanisms including raw context, compressed memory, spatial summaries, and state-space recurrence, the study finds that raw context significantly improves open-domain return performance, while block-wise state-space recurrence achieves the best results on this task, whereas compact compression tends to discard critical information.
Existing evaluation methods struggle to assess whether multimodal agents retain critical visual evidence necessary for reasoning, particularly lacking fine-grained and dynamic state-change evaluations. This work proposes MemEye, a novel framework that introduces visual centrality into multimodal memory assessment. It constructs a new benchmark spanning eight everyday scenarios, structured along two dimensions: granularity of visual evidence (from scene-level to pixel-level) and usage patterns (from single-evidence to temporally evolved synthesis). Through ablation-based verification gating, multi-granularity annotations, and cross-temporal state tracking, the study systematically evaluates 13 memory mechanisms and four vision-language models. Results reveal that current systems generally fail to preserve fine-grained details or reason about temporal state transitions, underscoring the critical roles of evidence routing, temporal tracking, and detail extraction in effective multimodal reasoning.
This work addresses the lack of direct evaluation of long-term memory content in large language model agents, which currently rely solely on downstream behavioral proxies that hinder auditability of retained user states. The authors propose treating long-term memory as an auditable artifact by directly assessing its quality through reconstruction of latent, structured user states. To this end, they introduce MEMPROBE—the first benchmark for memory recovery capability—featuring a synthetic ground-truth repository of hidden user states, simulated user trajectories, controlled information leakage tasks, and balanced state dimensions. Evaluations under both full-storage and top-k retrieval settings reveal a significant gap between task completion performance and memory fidelity: despite high task success rates across 50 users and 1,550 targets, memory recovery rates hover around 0.6 and further decline under top-k retrieval, underscoring a critical disconnect between functional assistance and faithful memory retention.
This work addresses a critical limitation in existing agent memory evaluation methods, which decouple memory from action and thus fail to capture how memory guides decision-making in real-world scenarios. To bridge this gap, the authors propose MemoryArena—a unified benchmark framework for evaluating agent memory across multiple sessions. MemoryArena introduces human-designed inter-task dependent subtasks—such as web navigation, preference-constrained planning, progressive retrieval, and sequential reasoning—that explicitly couple memory acquisition with action selection, establishing a novel paradigm for assessing memory in multi-session, task-dependent settings. Experimental results reveal that even agents excelling on current long-context benchmarks suffer significant performance degradation within MemoryArena, exposing a key blind spot in contemporary memory evaluation approaches.
Existing world models lack a unified open-domain closed-loop benchmark, making it difficult to systematically evaluate their memory consistency and action control capabilities. To address this gap, this work proposes MIND, a benchmark comprising 250 high-resolution (1080p/24 FPS) multi-view synchronized video sequences that span a diverse action space—including variations in movement speed and camera rotation—and introduces a closed-loop interactive evaluation framework. Additionally, we present MIND-World, the first Video-to-World baseline method designed for open-domain scenarios. Experimental results demonstrate that current models still face significant challenges in long-term memory stability and generalization across actions, highlighting MIND as a reliable platform for future research in world modeling.
This work addresses the limitations of traditional Embodied Question Answering (EQA), which relies on episodic evaluation and struggles to support real-world robots in continuous tasks requiring accumulation and reuse of prior knowledge. Focusing on continuous, multi-turn EQA scenarios, the study systematically investigates the impact of memory architectures on embodied agent performance and proposes a structured, spatially anchored visual memory mechanism that maps persistent observations into a metric 3D geometric space. The approach reveals, for the first time, fundamental bottlenecks in existing memory designs and demonstrates that 3D geometry–based spatial anchoring is crucial for overcoming the trade-off between answer accuracy and navigation efficiency. Experiments show that the proposed architecture significantly improves accuracy while reducing navigation cost in simulation, and its effectiveness in continuous intelligent interaction is further validated on a real mobile robot.
为解决长期互动任务中代理难以可靠保持环境信息的问题,本文通过引入EmbodiedMemory-Bench评估基准和Embodied-Memorizer记忆系统来提高代理的记忆能力。
This study addresses the limitation of existing agent memory evaluations that focus solely on final accuracy while overlooking the complex stability-plasticity trade-off. To this end, this work proposes MemProbe, a novel framework that pioneers cognitive experimental paradigms for memory diagnostics. By integrating interference and misinformation paradigms, it constructs a diagnostic suite comprising 56 scenarios to conduct a unified behavioral analysis of six incremental memory systems. The findings reveal that systems with comparable overall scores exhibit significant disparities in memory updating, retention, and attribution mechanisms. Ultimately, this research facilitates a paradigm shift from opaque black-box scoring toward transparent, interpretable behavioral profiling.
本文探讨了自回归视频生成中由于上下文窗口限制导致的历史信息丢失问题,并通过五种视角系统综述了记忆机制以维持时间持久性。
This study addresses the challenge of effectively transferring historical memory for embodied agents under varying environmental conditions. Leveraging a shared frozen vision-language model within a simulated warehouse environment, it establishes a benchmark to systematically evaluate the robustness of six memory representations and a working memory baseline against variations in starting poses, path availability, and historical relevance. This work provides the first quantification of the independent sensitivity of different memory representations to experiential mismatches, revealing that robustness along a single dimension does not guarantee overall generalization. Furthermore, it demonstrates that episodic memory is susceptible to interference while summary memory exhibits low retention during path blockages, thereby underscoring the necessity of jointly evaluating both the storage and utilization mechanisms of memory.