DashAct: A Progressive Diagnostic Benchmark for GUI Agents in Interactive Dashboard Analysis

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of fine-grained diagnostic capabilities for GUI agent failures in existing dashboard benchmarks. We construct the first interactive dashboard analysis diagnostic benchmark, comprising 357 human-verified trajectories. By integrating milestone dependency modeling with hierarchical goal annotation, we design a context-recovering progressive cascading evaluation mechanism that precisely localizes execution, prediction, and visual grounding bottlenecks within individual tasks. Experimental results reveal that current models exhibit limited performance even when provided with additional support, thereby offering clear guidance for optimizing GUI agents.
📝 Abstract
Interactive dashboards require users to reveal and connect evidence across stateful interactions. Although graphical user interface (GUI) agents could automate this process, existing dashboard benchmarks primarily report final answers or task success. They provide limited insight into whether failures arise from maintaining the analytical process, selecting actions, or grounding visual targets. We introduce DashAct, to our knowledge the first benchmark to diagnose these failures at a fine-grained level within the same dashboard task. DashAct contains 357 human-verified interaction trajectories with milestone dependencies and hierarchical target annotations. Its progressive diagnostic cascade evaluates end-to-end execution, restores verified context for next-action prediction, and provides target semantics and a local view for visual grounding. By progressively restoring the conditions for success, DashAct measures the minimum support an agent needs to recover rather than scoring isolated skills. Experiments show that current models struggle even as support is added. The cascade outcomes reveal bottlenecks hidden by end-to-end scores and provide actionable guidance for improving GUI agents.
Problem

Research questions and friction points this paper is trying to address.

GUI agents
interactive dashboard
benchmark
failure diagnosis
visual grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

GUI Agents
Diagnostic Benchmark
Progressive Cascade
Visual Grounding
Interactive Dashboards
🔎 Similar Papers