🤖 AI Summary
This study addresses the computational waste and loss of critical details in existing agent memory systems caused by query-agnostic preprocessing. To this end, we propose a reinforcement learning-based, on-demand multimodal memory curation framework. Methodologically, the framework dynamically schedules heterogeneous models to enable fine-grained runtime resource allocation. Furthermore, we introduce goal-decoupled advantage estimation and prefix marginal utility evaluation mechanisms to resolve fine-grained credit assignment and multi-objective collaborative optimization challenges in multi-step decision-making. Experimental results demonstrate that the proposed approach significantly expands the Pareto frontier of performance, cost, and latency across five benchmarks, effectively accommodating diverse optimization preferences.
📝 Abstract
Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can incur unnecessary preprocessing cost and discard details that later prove essential. Recent studies have begun shifting memory processing toward runtime adaptation, but typically specialize in particular operations or fixed processing schemes, leaving flexible control over performance, cost, and latency largely underexplored. To address this challenge, we present \textbf{MemPilot}, a flexible framework that orchestrates on-demand memory curation under different performance--cost--latency preferences. Specifically, we optimize a multi-step LLM policy via reinforcement learning to iteratively choose between retrieving from query-agnostic memory and delegating query-specific curation of raw multimodal history to heterogeneous LLMs and VLMs. The policy jointly controls evidence amount, curation instructions, model selection, and visual access, enabling fine-grained allocation of runtime computation. To optimize this policy under competing objectives, we adapt objective-wise advantage decoupling by separately estimating each objective's advantage before aggregation. Moreover, we introduce prefix-based marginal utility estimation for fine-grained credit assignment across multi-step rollouts. Experiments on five multimodal agent-memory benchmarks demonstrate favorable performance--cost--latency trade-offs across optimization preferences, with preference sweeps yielding broader frontiers than existing trade-off-aware baselines.