๐ค AI Summary
Existing benchmarks for long-term memory evaluation primarily focus on factual recall and simple retrieval, failing to adequately assess large language modelsโ (LLMsโ) capacity to organize and leverage complex memory structures. To address this gap, this work introduces StructMemEval, a novel benchmark that systematically evaluates LLM agentsโ ability to construct and utilize structured long-term memory in complex reasoning scenarios. StructMemEval emphasizes structural organization through tasks such as transaction ledgers, to-do lists, and tree-structured data. Experimental results demonstrate that mainstream LLMs struggle to autonomously organize memories into coherent structures without explicit guidance, whereas memory-augmented agents equipped with structural prompts achieve significantly higher task success rates. This benchmark thus fills a critical void in the evaluation of sophisticated memory architectures for LLM-based agents.
๐ Abstract
Modern LLM-based agents and chat assistants rely on long-term memory frameworks to store reusable knowledge, recall user preferences, and augment reasoning. As researchers create more complex memory architectures, it becomes increasingly difficult to analyze their capabilities and guide future memory designs. Most long-term memory benchmarks focus on simple fact retention, multi-hop recall, and time-based changes. While undoubtedly important, these capabilities can often be achieved with simple retrieval-augmented LLMs and do not test complex memory hierarchies. To bridge this gap, we propose StructMemEval - a benchmark that tests the agent's ability to organize its long-term memory, not just factual recall. We gather a suite of tasks that humans solve by organizing their knowledge in a specific structure: transaction ledgers, to-do lists, trees and others. Our initial experiments show that simple retrieval-augmented LLMs struggle with these tasks, whereas memory agents can reliably solve them if prompted how to organize their memory. However, we also find that modern LLMs do not always recognize the memory structure when not prompted to do so. This highlights an important direction for future improvements in both LLM training and memory frameworks.