🤖 AI Summary
This study addresses the challenges of storage location conflicts—between model weights and context—and performance instability caused by holistic trajectory mixing in self-improving GUI agents. To overcome these limitations, we propose a fine-grained component routing mechanism based on experience attributes. Specifically, this method decomposes experiences into four component types, including locators and programs, and dynamically routes each to its optimal storage location according to recurrence and state-conditionality rules. As the first component-level routing approach, this work reveals optimal storage strategies for distinct components and establishes generalizable routing principles. Experimental results demonstrate that the proposed method outperforms whole-trajectory baselines by an average of 3.5 points across multiple benchmarks, while the derived routing rules exhibit strong generalization to unseen models.
📝 Abstract
Self-improving GUI agents keep the trajectories they produce and return them to the agent, by fine-tuning or by retrieval into the prompt, and studies that compare the two destinations disagree. We attribute this to the unit of experience: a trajectory bundles items with different properties, so a conclusion about the bundle depends on its mix. To address this, (i) we introduce component routing, which splits the experience into locators, procedures, state facts and lessons and sends each component to the context or to the weights, compared on the same items across three backbone families, two environments and three seeds. One pool has two destinations: locators and lessons win in the weights, procedures and state facts in the context. (ii) We fit a rule in two properties measured before any training, recurrence and state-conditionality; it recovers the destination of a held-out backbone family in 24 of 24 cells, two interventions move a component toward the boundary, and routing by the rule beats every whole-trajectory baseline and, by +3.5 points on average, the better single destination of each backbone. (iii) We identify how training and producer-consumer differences change the value of the two destinations: note readout decreases after the same component is written into the weights, most for the items that recur most, context gains increase with the information gap, and weights gains decrease with the policy gap. Code and data will be released.