🤖 AI Summary
This study addresses the limitation of existing memory systems that compress documents during write operations, which hinders accurate responses to queries involving historical versions of organizational decisions. To overcome this, we propose Mem++, a non-destructive memory framework that replaces write-time distillation with read-time selection. Mem++ preserves all document versions intact and delegates version selection to the answering model through time-aware retrieval and hybrid lexical-semantic ranking. Built upon a large language model (LLM) agent architecture, this framework effectively tackles long-term organizational question answering. Experimental results demonstrate that Mem++ outperforms the strongest baseline by 8.0 to 13.1 points on the OrgMemBench benchmark, significantly surpassing conventional retrieval-augmented generation (RAG) approaches and mainstream long-term memory systems.
📝 Abstract
Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time. However, most memory systems compress the record at write time. By distilling each document into facts, notes or graph edges, these methods fix what can be answered before any question is asked. To address this, we propose Mem++, a non-destructive memory framework shifting from write-time distillation to read-time selection. Mem++ stores every document whole with its date and author, and it calls no generative model at write time. At read time, it retrieves only documents dated up to the time a question asks about and fuses lexical and semantic rankings. Unlike systems that overwrite older versions, Mem++ keeps them and leaves the choice to the answering model. Evaluations on the organizational benchmark OrgMemBench demonstrate that Mem++ surpasses the strongest memory system baseline by 8.0 to 13.1 points across two answering models. With gpt-4.1-mini, it also achieves the best overall score, 2.6 points above RAG. In addition, Mem++ achieves the best average LLM-judge score on LoCoMo and ranks second on LongMemEval-S, behind only its entity-graph variant. Code for benchmark evaluation is available at https://github.com/AIDAChip-Inc/mem-plus-plus.