๐ค AI Summary
This study addresses the limitation of existing agent memory frameworks in counterfactual exploration, which hinders the low-cost acquisition of alternative action feedback for decision optimization. We propose a reinforcement learning-based counterfactual memory system that leverages an executable world model to validate alternative actions and store verified experiences. By integrating a frozen foundation model with an independent memory policy, the system enables efficient cross-task reuse while balancing success rates against interaction costs through learned strategies. Evaluated across 12 benchmarks spanning six domains, our approach yields an average performance improvement of 12.6 percentage points and significantly reduces token consumption by 7.7%โ42.0%. These results demonstrate the synergistic optimization of both agent decision-making efficacy and resource efficiency.
๐ Abstract
Existing agent memory frameworks mainly create memory through an agent's interaction with the factual world, e.g., remembering feedback from actions taken to improve performance on future tasks. However, these frameworks seldom ask the"what if"question during memory construction: what if a different action had been taken, would the feedback have changed, and how could this feedback become useful memory? Obtaining such feedback directly in an active environment can be expensive and can alter the state needed for comparison. In this work, we introduce COUNTERMEM, a reinforcement-learning framework for constructing and using verified counterfactual memory across tasks. After a failed action, COUNTERMEM evaluates local alternatives from a copy or reset of the original state using executable world models, such as tests, proof checkers, and solvers. It stores improvements with the original and corrected actions, checked outcomes, and conditions for reuse. A learned memory-use policy selects a retrieved record or skips memory to balance task success and interaction cost, while the base LLM remains fixed. Both memory and policy are frozen during held-out evaluation. We evaluate COUNTERMEM on 12 benchmark settings across six domains. With gpt-oss-120b, COUNTERMEM improves both ReAct and Reflexion on all 12 benchmarks across six domains, averaging a gain of 12.6 percentage points over their unaugmented versions. In the four-domain comparison across two backbones, task-run tokens decrease by 7.7-42.0%, excluding offline selector-training costs. Further analyses show that removing verification or persistent storage weakens the gains, while applying verified corrections to unsuitable decisions can reverse them. Code will be released upon acceptance.