Learning What to Remember: Long-horizon Counterfactual Memory Optimization

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the credit assignment challenge in long-term memory for large language models by proposing the MGPO algorithm. This method introduces a novel incremental value-based credit assignment mechanism for memory rewriting, which isolates the marginal contribution of individual rewrites through long-horizon counterfactual evaluation, thereby converting delayed utility into direct learning signals. By integrating document-level structured supervision with reinforcement learning, MGPO optimizes information persistence strategies. Experimental results demonstrate that MGPO substantially enhances information extraction performance while reducing average memory length by nearly 80%. Furthermore, it enables cross-domain reuse of memory strategies and facilitates downstream transfer without fine-tuning.
📝 Abstract
Persistent textual memory allows language models to carry information across long interactions, but learning what to remember is fundamentally a credit-assignment problem. A memory rewrite may only become useful many steps later, while much of the observed utility may be inherited from information already stored before the rewrite. We introduce Memory Gain Policy Optimization (MGPO), which isolates the incremental value of each memory rewrite by crediting it for its marginal contribution to current and future downstream utility. This turns delayed memory utility into a direct learning signal for optimizing what information should persist. We study MGPO on document-level information extraction, where structured supervision makes the effects of individual memory updates directly measurable. MGPO improves extraction while reducing average memory length by nearly 80% relative to the initial memory policy before optimization. The learned memory policy also supports reuse and transfer across domains, downstream models without further training. These results show that effective memory learning depends not only on preserving useful information, but on identifying which memory updates create lasting incremental value.
Problem

Research questions and friction points this paper is trying to address.

credit assignment
long-horizon memory
memory optimization
language models
delayed utility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Memory Gain Policy Optimization
Credit Assignment
Counterfactual Memory
Long-horizon
Policy Optimization