DyadMem: A Long-Term Memory Benchmark of How Agents Work with Users

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing agent benchmarks, which overlook user-agent relational memory and rely solely on final question-answering (QA) evaluations, thereby lacking full-pipeline reliability. To this end, this work proposes the DyadMem benchmark and the URAM framework. By jointly annotating user facts and six categories of relational memory across multi-session trajectories, it constructs a dual QA evaluation system comprising Gold-Memory and Full-Pipeline assessments. This establishes the first dual-domain, full-pipeline evaluation paradigm that explicitly models the capture, update, and retrieval processes of memory. Experimental results demonstrate that state-of-the-art models exhibit significant performance degradation under full-pipeline evaluation, revealing fine-grained deficiencies such as low recall rates and insecure deletion. These findings validate the effectiveness of the proposed approach in rigorously assessing relational memory capabilities in AI agents.
📝 Abstract
Long-term agents must remember not only what is true about a user, but also how a particular agent should work with that user as their shared history evolves. Existing benchmarks primarily supervise user facts and preferences or experience reusable across users, leaving this relationship-specific agent memory implicit. Additionally, most prior works measure the model solely with final-answer QA over long interaction histories, making the assessment still incomplete and unreliable. To this end, we introduce DyadMem with the proposed new definition User-conditioned Relational Agent Memory (URAM). DyadMem jointly annotates user-side memory and URAM along the same multi-session trajectories, resulting in 6 memory categories. To summarize, it includes 3,065 episodes, 50,961 sessions, and 61,210 QA instances, with extensive session-level Capture and Update gold annotations, query-level Recall support, and two QA settings: Gold-Memory and Full-Pipeline. Across 16 open-weight and 4 proprietary models, Gold-Memory QA is consistently strong, yet Full-Pipeline QA drops sharply. Such a gap explicitly supports our fine-grained evaluation design. Additionally, several quantitative results further reveal low Capture recall, incomplete Recall, and unsafe-deletion issues arising from even the frontier LLMs. We further conduct a rigorous experiment to validate the effectiveness of our URAM and observe the positive effects for all 20 models. In summary, DyadMem is a dual-domain, full-pipeline memory benchmark with extensive annotation efforts for advancing the domain's development.
Problem

Research questions and friction points this paper is trying to address.

long-term memory
relational agent memory
benchmark evaluation
memory management
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

User-conditioned Relational Agent Memory
Long-term memory benchmark
Fine-grained evaluation
Full-pipeline QA
Memory annotation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yifei Tao
Nanyang Technological University
X
Xinyu Zhong
Stepfun
Henry Hengyuan Zhao
Henry Hengyuan Zhao
Ph.D. student at National University of Singapore
Multimodal ReasoningAI AgentHuman-AI Interaction
F
Fanyi Wang
Stepfun
T
Tengda Guo
Stepfun
W
Wentao Qiu
Stepfun
Y
Ying Wang
University of Hong Kong
L
Liujian Tang
Stepfun