🤖 AI Summary
This study addresses the limitation of existing LLM memory systems that rely on passive repair, where retrieval failures already incur tangible interaction losses. To this end, we propose MemDream, a framework introducing a self-probing memory evolution paradigm that proactively probes, diagnoses, and repairs memory graphs through offline "dreaming" cycles. Methodologically, MemDream employs a multi-agent collaborative architecture comprising Dreamer, Analyst, and Consolidator agents, optimized via Group Relative Policy Optimization (GRPO) to refine repair strategies. Furthermore, it incorporates reversible forgetting and soft decay mechanisms to enable preventive maintenance and self-evolution of memory. Evaluated on the LoCoMo and MemoryAgentBench benchmarks, MemDream significantly outperforms existing baselines, achieving a 4.5-point improvement in F1 score and a 9.1-point increase in overall composite score.
📝 Abstract
Memory is essential for enabling LLM-based agents to maintain coherent, personalized behavior over long-horizon interactions. However, existing memory systems share a fundamental limitation: they never proactively test their own memory, repairing it only after real queries expose weaknesses. This reactive paradigm means every retrieval failure corresponds to a real interaction in which the cost has already been paid. We propose MemDream, a framework that enables self-probing memory evolution for LLM agents. Our framework periodically enters offline dream cycles where three specialized agents (Dreamer, Analyst, Consolidator) collaboratively probe, diagnose, and repair the memory graph before failures occur. A policy trained via Group Relative Policy Optimization learns which repair operations produce durable retrieval improvements, while a soft decay mechanism provides reversible forgetting driven by the same anticipatory signal. Experiments on LoCoMo and MemoryAgentBench demonstrate that MemDream improves answer F1 by 4.5 points on LoCoMo and achieves a 9.1-point higher overall score on MAB over the strongest reactive-evolution baselines.