🤖 AI Summary
This study addresses the challenges of persistent cross-session memory, dynamic updating, and efficient retrieval in AI agents by proposing a multi-level memory architecture grounded in explicit provenance tracing. The proposed method organically integrates episodic records, semantic relational graphs, community summaries, and persistent graph states through source linking, thereby enabling incremental updates, precise deletion, and full-lifecycle management of memory structures. Experimental evaluations demonstrate that the system substantially reduces token consumption across multiple benchmarks while effectively balancing high accuracy with computational efficiency. Consequently, this work provides a scalable new paradigm for long-horizon agent memory management.
📝 Abstract
AI agents require memory that preserves information across user-agent exchanges, user-to-user conversations, and group conversations with or without agent participation, while supporting updates as evidence changes or is removed. We present Stashbird, an agent memory system that links source episodes to derived memory state through explicit provenance. Stashbird organizes memory into episodic records, semantic relations, community summaries, and persisted graph state, with lifecycle operations for incremental updates and episode-level deletion. We evaluate question-answering accuracy and model-facing workload across four long-term memory benchmarks. On LoCoMo, Stashbird uses 76.4x fewer ingestion prompt tokens than Graphiti. Compared with reproduced Hindsight on the same benchmark, it uses 8.1x fewer retrieval prompt tokens, with accuracy 1.6 percentage points lower. It achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.