🤖 AI Summary
This study addresses the perceptual fragility of recommendation agents in noisy, heterogeneous web pages and their inefficiency in long-context reasoning by proposing a platform-agnostic framework for human-like perception and dynamic memory. Methodologically, it combines OCR with multimodal large language models to extract structured information from screenshots. A chunked sequential update strategy is designed to maintain a fixed-length interaction history, enabling linear-complexity long-term preference modeling. Furthermore, a multi-memory GRPO reinforcement learning variant is introduced to optimize training. Experimental results demonstrate that the proposed approach outperforms state-of-the-art baselines by an average of 5.16% across search, ranking, and judgment tasks on three datasets.
📝 Abstract
Recent Recommendation Agents (RecAgents) offer a promising alternative by shifting recommendation to an active, user-side paradigm, where generative agents autonomously perceive external platforms, reason over user preferences, and execute decisions. However, existing RecAgents still suffer from two critical limitations: brittle item perception based on noisy and heterogeneous item pages, and inefficient long-context reasoning over extended user histories and multi-step interaction traces. To address these challenges, we propose a novel recommendation agent framework, termed as ReMem, that combines OCR-based multimodal perception with time-evolving dynamic memory. Instead of parsing raw HTML, ReMem observes item pages through screenshots and extracts structured multimodal information via an OCR tool, enabling a more humanoid and platform-agnostic perception mechanism. To support long-horizon preference modeling, ReMem further introduces a chunk-wise sequential memory update strategy, where the agent selectively maintains a fixed-size memory of informative historical interactions while processing arbitrarily long contexts with linear inference complexity and bounded context length. This design allows the agent to preserve evolving user preferences without relying on external memory modules or disrupting the standard autoregressive generation process. To enhance the dynamic memory instruction, we further develop a multi-memory GRPO variant, which propagates the final-answer advantage to all intermediate conversations that contribute to the final response. Extensive experiments on three datasets demonstrate that ReMem consistently outperforms state-of-the-art baselines, achieving an average improvement of 5.16\% across three recommendation agent tasks, namely searching, ranking, and judging.