🤖 AI Summary
This study addresses the challenges of sparse evidence retrieval in long conversational histories and the inefficiency caused by fixed granularity in personalized large language models. To this end, it proposes a four-layer progressive disclosure memory architecture. This approach pioneers an agent-controlled dynamic retrieval mechanism that replaces conventional flat storage paradigms. By coordinating a controller for on-demand deep retrieval with a note synthesizer, the architecture achieves an adaptive cost-fidelity trade-off, enabling early stopping for simple queries and deeper exploration for complex ones. Evaluated on the LongMemEval benchmark, the proposed method demonstrates robust long-context memory reasoning performance while accessing merely 8% of the conversational data, highlighting its effectiveness in balancing computational efficiency with retrieval accuracy.
📝 Abstract
Personalized LLM assistants must recover sparse evidence from long conversation histories across queries of varying complexity. We introduce APDMem (Agent-controlled Progressive Disclosure Memory), a hierarchical long-term memory architecture that applies progressive disclosure to memory retrieval. Rather than relying on a flat memory store or fixed retrieval granularity, APDMem represents conversation history as four progressively detailed layers: thematic summaries, personalized key facts, turn-level evidence notes, and raw messages. At inference time, a controller applies progressive disclosure to the memory hierarchy: it first reads high-level summaries and drills into finer evidence only when needed. This creates an adaptive cost-fidelity trade-off: simple queries can terminate early, while complex temporal, multi-hop, or exact-evidence queries trigger deeper inspection. A note synthesizer converts retrieved evidence into a query-focused structure that consolidates facts, orders events, and flags contradictions before final answer generation. Experiments on LongMemEval show that APDMem achieves strong performance for long-context memory reasoning while accessing only 8% of the total conversations.