🤖 AI Summary
This study addresses the limited cross-task generalization of existing memory methods for LLM-based agents, which typically rely on a single representation. We formulate memory management as a routing problem and propose a heterogeneous memory routing paradigm coupled with a staged supervised synthesis pipeline. By incorporating content-aware probing, short-term memory gating, and selective multi-provider storage, our approach jointly optimizes retrieval, injection, and storage operations. This method overcomes the inherent bottleneck of single-representation memory systems, achieving an average accuracy improvement of 10.0% while significantly reducing execution steps. Notably, the routing overhead remains below 0.3%. Overall, the proposed framework comprehensively outperforms existing single-memory baselines, demonstrating both effectiveness and efficiency in agent memory management.
📝 Abstract
Current agents remain largely stateless across tasks, limiting their ability to continually improve from prior interactions and making memory essential for long-horizon agentic behavior. Existing memory methods seek to reuse past experience, but most rely on a single memory representation (e.g., trajectories, reflections, skills, structured knowledge) whose effectiveness varies across task distributions. Rethinking this design space, we evaluate 13 memory methods and find that no single method generalizes across benchmarks, revealing the potential of managing heterogeneous memory providers. We formulate agent memory as a routing problem in which a memory agent decides which memory provider to retrieve from, whether to inject short-term memory, and which providers should store the resulting experience. Based on this perspective, we propose MemAgent, featuring a content-aware routing architecture and a training-data synthesis pipeline. The routing architecture combines content-aware probing before retrieval, short-term memory gating during execution, and selective multi-provider storage, while the training pipeline synthesizes phase-specific supervision for routing decisions. Across GAIA, WebWalkerQA, and xBench-DS, MemAgent improves average accuracy by 10.0% and outperforms every individual memory method across all three benchmarks. These gains come with less than 0.3% routing overhead and a 12% reduction in average task steps.