🤖 AI Summary
This study addresses the challenge of horizontally evaluating long-term memory systems for LLM-based agents, which is hindered by heterogeneous representation and retrieval mechanisms. To this end, it proposes a unified, controlled evaluation framework grounded in 5W (Who, What, When, Where, Why) conversational memory. Methodologically, an AdaptiveGraph structure integrating temporal edges with personalized PageRank diffusion is constructed to enable fair comparisons across multiple strategies, including graph traversal, adaptive graph diffusion, and a BM25 lexical baseline. The findings reveal that a strong lexical baseline outperforms graph-based methods under most configurations. Furthermore, information extraction quality and vocabulary normalization significantly impact graph retrieval performance. These insights provide critical empirical guidance for designing agent memory systems.
📝 Abstract
Long-term memory lets LLM agents recall past interactions and remain consistent across sessions, but memory systems are hard to compare because they often vary in representation, indexing, retrieval, and evaluation. We present a controlled evaluation framework based on shared 5W-style conversational memories. Localized graph configurations traverse a common base graph; AdaptiveGraph adds chronological edges and Personalized PageRank diffusion. We also evaluate BM25 over the same extracted notes and OpenClaw as a raw-input external reference. Retrieval rankings vary across memory settings. On LongMemEval-S, AdaptiveGraph is the strongest graph configuration at 0.844 MRR, but BM25 reaches 0.867 and OpenClaw 0.880. On ATANT Core, localized graph traversal outperforms diffusion and BM25, whereas BM25 leads the stress rounds. Reducing LongMemEval-S within the tested range does not reproduce the ATANT diffusion penalty, but the smallest tested store remains larger than ATANT Core, so store size cannot be ruled out. The penalty also persists under a permissive content-match criterion. Vocabulary normalization and extraction quality substantially affect graph retrieval, and missing extraction tags are common among top-five misses. Retrieval strategies should therefore be evaluated jointly with the memory setting and against strong lexical baselines.