Evaluating Memory Structure in LLM Agents

๐Ÿ“… 2026-02-11
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing benchmarks for long-term memory evaluation primarily focus on factual recall and simple retrieval, failing to adequately assess large language modelsโ€™ (LLMsโ€™) capacity to organize and leverage complex memory structures. To address this gap, this work introduces StructMemEval, a novel benchmark that systematically evaluates LLM agentsโ€™ ability to construct and utilize structured long-term memory in complex reasoning scenarios. StructMemEval emphasizes structural organization through tasks such as transaction ledgers, to-do lists, and tree-structured data. Experimental results demonstrate that mainstream LLMs struggle to autonomously organize memories into coherent structures without explicit guidance, whereas memory-augmented agents equipped with structural prompts achieve significantly higher task success rates. This benchmark thus fills a critical void in the evaluation of sophisticated memory architectures for LLM-based agents.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsCognitive Modeling & Cognitive Systems: Agent Architectures

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
๐Ÿ“ Abstract
Modern LLM-based agents and chat assistants rely on long-term memory frameworks to store reusable knowledge, recall user preferences, and augment reasoning. As researchers create more complex memory architectures, it becomes increasingly difficult to analyze their capabilities and guide future memory designs. Most long-term memory benchmarks focus on simple fact retention, multi-hop recall, and time-based changes. While undoubtedly important, these capabilities can often be achieved with simple retrieval-augmented LLMs and do not test complex memory hierarchies. To bridge this gap, we propose StructMemEval - a benchmark that tests the agent's ability to organize its long-term memory, not just factual recall. We gather a suite of tasks that humans solve by organizing their knowledge in a specific structure: transaction ledgers, to-do lists, trees and others. Our initial experiments show that simple retrieval-augmented LLMs struggle with these tasks, whereas memory agents can reliably solve them if prompted how to organize their memory. However, we also find that modern LLMs do not always recognize the memory structure when not prompted to do so. This highlights an important direction for future improvements in both LLM training and memory frameworks.
Problem

Research questions and friction points this paper is trying to address.

memory structure
LLM agents
long-term memory
memory organization
benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

structured memory
LLM agents
memory organization
benchmarking
retrieval-augmented generation
A
Alina Shutova
HSE University
A
Alexandra Olenina
Yandex
I
Ivan Vinogradov
YSDA
A
Anton Sinitsin
Yandex