🤖 AI Summary
This study addresses the challenge that memory retrieval for LLM agents relies on costly execution feedback, hindering efficient learning of optimal memory sets. To this end, this work proposes a closed-loop Gaussian probing strategy based on Expected Value of Sample Information (EVSI), which computes set-level execution gains to guide directed exploration of surrogate memory sets. By integrating a frozen executor with a shared scorer, the approach optimizes the allocation of limited training resources and enhances feedback coverage. Consequently, this method enables efficient memory selection without requiring test-time probing. Experiments on benchmarks such as ALFWorld demonstrate state-of-the-art success rates, significantly improving both memory decision quality and feedback utilization efficiency.
📝 Abstract
Large language model (LLM) agents reuse external memory to guide new tasks, but effective retrieval requires learning which memory sets improve execution. Such learning relies on costly outcome feedback: ordinary retrieval observes only executed sets, while evaluating alternatives requires additional rollouts. We introduce \textsc{UpliftMem}, which learns memory retrieval from set-level execution uplift relative to the same executor without memory. A theoretical analysis of how retrieval preferences restrict feedback coverage motivates targeted probing of alternative memory sets. Probe selection follows an expected value of sample information (EVSI) criterion, derived in closed form under a correlated Gaussian model, to allocate limited training rollouts according to their expected improvement in local retrieval decisions. The shared scorer is trained with a frozen executor and selects memory sets without test-time probes. Across ALFWorld, WebShop, and BigCodeBench, \textsc{UpliftMem} achieves the best success rates among evaluated baselines on the main evaluation sets. Controlled fixed-store and matched probe budget evaluations further demonstrate improved memory-use decisions and more effective use of execution feedback.