How Can Recommendation Feedback Evolve Agent Memory?

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of agent memory evolution in recommender systems caused by delayed feedback, noise, and attribution difficulties. To tackle these challenges, we propose TIDE, a framework that employs a trajectory-guided memory evolution mechanism. By integrating a capacity-limited experience pool with reinforcement, crossover, mutation, and elimination operators, TIDE effectively optimizes the memory population. Furthermore, this work introduces the MEG metric to quantify future task utility gains, enabling precise multi-memory attribution and adaptive updates through temporal and semantic credit assignment. Experimental results demonstrate that TIDE achieves a 7.75% improvement in offline MEG, yields significant gains in click-through rate and activation rate in online A/B testing, and attains the lowest error on standard benchmarks.
📝 Abstract
Content-generation agents continuously receive impressions, clicks, conversions, and negative feedback from recommendation systems, providing real-world outcome signals for memory evolution. However, these signals are delayed and noisy, confounded by audience composition, placement, and recommendation policies, and may result from the combined influence of multiple memories, making accurate attribution difficult. Existing methods rely primarily on immediate feedback or semantic retrieval and therefore struggle to reliably translate recommendation outcomes into memory fitness. To address this challenge, we propose TIDE (Trajectory-Informed Directed Memory Evolution), an external memory evolution framework driven by delayed recommendation feedback. We further introduce Memory Evolution Gain (MEG), which measures the utility improvement of evolved memory over a no memory baseline on strictly future tasks. TIDE treats memory as a capacity-constrained population of experiences: temporal and semantic credit assignment estimates contextual fitness, while responsibility credit distributes outcome signals according to the memories referenced during generation. These signals are then used to reinforce, crossover, mutate, or evict memories. On an e-commerce membership marketing content-generation agent, TIDE achieves a +7.75-percentage-point MEG in offline temporal replay and significantly improves both unique click-through rate (UCTR) and activation rate in an online A/B test. On a delayed-label benchmark, TIDE achieves the lowest mean absolute error (MAE) and root mean squared error (RMSE) and the highest MEG among the compared methods, demonstrating its effectiveness.
Problem

Research questions and friction points this paper is trying to address.

recommendation feedback
memory evolution
delayed feedback
credit assignment
content-generation agent
Innovation

Methods, ideas, or system contributions that make the work stand out.

Memory Evolution
Delayed Feedback
Credit Assignment
Content-Generation Agent
Recommendation Systems
S
Shanwen Mao
Harbin Institute of Technology, Harbin, China
Mingming Li
Mingming Li
Zhejiang University
FabricationHuman-Computer Interaction
H
Hao Zhang
Harbin Institute of Technology, Harbin, China
Z
Zhiheng Li
Institute of Automation, Chinese Academy of Sciences, Beijing, China
Y
Yige Wang
Alibaba Group, Hangzhou, China
P
Penghua Yu
Alibaba Group, Hangzhou, China
J
Junxiong Zhu
Alibaba Group, Hangzhou, China