GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of intent alignment, fault diagnosis, and pattern reuse in the automatic optimization of GUI agent execution frameworks. To this end, we propose an evidence-driven framework evolution mechanism. Specifically, our method localizes behavioral discrepancies through visual state alignment and aggregates cross-task failure patterns by combining repeated execution with joint evidence analysis. It subsequently generates constrained source code edits and performs predictive verification, thereby enabling self-improvement and generalization while keeping the underlying model frozen. Extensive experiments on benchmarks such as OSWorld demonstrate that the proposed approach significantly outperforms existing baselines, yielding a performance gain of 12.33 points for Qwen3-VL.
📝 Abstract
The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled. Compared with harness optimization for non-GUI agents, automatically optimizing this harness poses three coupled challenges: reconciling model intent with observed visual effects, diagnosing failures under variable execution outcomes, and identifying recurrent failure patterns across tasks and translating them into reusable runtime changes. We introduce GUI-HARVEST, an automatic harness optimizer that enables self-improving GUI agents with frozen backbone models. First, to ground diagnosis in observed action effects, it aligns model outputs and executed actions with before-and-after screenshots, tying findings to specific interface transitions. Second, to account for execution variability, it treats repeated runs of the same task as a joint evidence unit, using within-task comparisons to locate outcome-relevant behavioral differences. Third, it consolidates verified findings across tasks into recurring failure patterns, maps them to bounded source-code edits with predictions recorded before evaluation, and checks the predicted behavioral effects alongside task performance through repeated execution. Experiments on OSWorld-Verified show consistent held-out gains across six general-purpose open, GUI-specialized open, and proprietary backbone models; Qwen3-VL-32B-Instruct gains 12.33 points on the full suite. Frozen-harness transfer improves GPT-5 by 13.87 percentage points on WindowsAgentArena at 50 steps without further optimization. With the same backbone and initial harness, GUI-HARVEST outperforms Self-Harness and Meta-Harness, suggesting that GUI-specific diagnosis and validation help harness improvements generalize to unseen tasks. The code is available at https://github.com/GaryYang12345/GUI-HARVEST.
Problem

Research questions and friction points this paper is trying to address.

GUI agents
harness optimization
self-improving agents
failure diagnosis
evidence-driven
Innovation

Methods, ideas, or system contributions that make the work stand out.

GUI agents
harness optimization
self-improving
evidence-driven diagnosis
failure pattern generalization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Geyi Yang
The Chinese University of Hong Kong, Shenzhen
Z
Zikun Qu
The Chinese University of Hong Kong, Shenzhen
X
Xiang Li
Tianjin University
Zhiyong Wang
Zhiyong Wang
Harbin Institute of Technology, Shenzhen
Human-computer interaction
Min Zhang
Min Zhang
East China Normal University
S
Shipei Zeng
Shenzhen Research Institute of Big Data
Zhongxiang Dai
Zhongxiang Dai
Assistant Professor, The Chinese University of Hong Kong, Shenzhen
Machine LearningData-Centric AILarge Language ModelsMulti-Armed BanditsBayesian Optimization