DeltaReplay: Task-Relative Memory Reuse for Mobile GUI Agents

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of memory reuse in mobile GUI agents, where stored trajectories often imperfectly match new tasks. We propose a step-level memory reuse framework that models historical trajectories as transition graph paths and decouples actions from their parameters. Fine-grained retrieval is achieved based on page consistency and action generality. During execution, the system dynamically compares recorded steps with the current screen state to adaptively follow them, substitute parameters, or delegate control to a base agent, all without modifying the original memories. Experimental results demonstrate that this framework improves task success rates by 10.3 and 25.0 percentage points over baselines on AndroidWorld and SPA-Bench, respectively.
📝 Abstract
Memory-augmented mobile GUI agents store successful execution trajectories and reuse them in later tasks, but a stored trajectory rarely matches a new task exactly. The new task may use different parameters, share only some of its steps with a stored trajectory, or have no relevant record in memory. Forcing the agent to use irrelevant memory can mislead it, whereas discarding memory that may still be useful deprives it of guidance from past experience. To address this dilemma, we propose DeltaReplay, a step-level memory reuse framework that decides how to use existing memory without modifying it. We observe that the reusable part of a stored record is determined not by the record itself but by its relation to the new task, mainly through two factors: page-level consistency and action-level generality. We therefore store execution trajectories as paths in a transition graph, whose nodes (pages) and edges (actions between pages) capture these two factors. At reuse time, the action on each edge is split into a task-independent operation and task-specific parameters. DeltaReplay then compares each recorded step with the new task and the current screen, and decides whether to follow it, execute it after replacing its parameters, or leave it to the base agent. On AndroidWorld and SPA-Bench, DeltaReplay improves the task success rate over a base agent with the same backbone by up to 10.3 and 25.0 percentage points, respectively. These results indicate that deciding at each step how to use retrieved memory lets agents benefit even from partially matching trajectories.
Problem

Research questions and friction points this paper is trying to address.

Mobile GUI Agents
Memory Reuse
Trajectory Matching
Task-Relative Memory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Memory Reuse
Mobile GUI Agents
Transition Graph
Step-level Replay
Action Decoupling
🔎 Similar Papers
Y
Yudong Bai
Tsinghua University
Y
Yihong Chen
Tsinghua University
Quanming Yao
Quanming Yao
Associate Professor, EE Department, Tsinghua University
Machine Learning
Y
Yaqing Wang
Tsinghua University