🤖 AI Summary
This study addresses a critical limitation in existing GUI agent memory mechanisms, which rely solely on observational similarity to merge pages and consequently risk erroneously conflating states with distinct behaviors. To overcome this, we propose a training-free memory consolidation rule grounded in action-conditioned bisimulation. Departing from observational similarity, our approach leverages empirically predicted state graphs, adopting the consistency of action consequences and successor blocks as the sole criterion for merging, thereby enabling precise state abstraction. This method serves as a direct replacement for existing mechanisms. Closed-loop evaluations on the MiniWoB++ benchmark demonstrate that the proposed approach significantly improves task success rates, outperforming both memory-less baselines and traditional consolidation strategies. These results validate its effectiveness in enhancing agent memory management.
📝 Abstract
An agent that remembers what it did on a web page must decide when two pages count as the same. Memories built on observation similarity merge pages that look alike but behave differently, and GUIs are full of such pages: two tabs of one widget or two rows of one menu answer the same click differently. We define the merge rule as an action-conditioned bisimulation over the empirical predictive state graph a frozen agent fills as it acts. Two states merge only when their shared actions lead to agreeing outcomes and successor blocks under an affordance label. Observation similarity never enters the rule, and nothing is trained. It replaces the merge rule of an existing outcome-value memory, so a closed-loop comparison isolates it. On MiniWoB++ it raises success rate over a memoryless agent, while a control taking identical exploratory detours, the prior successor-representation merge, and the same criterion without action conditioning change nothing.