Score
Designs and implements automation scripts and tools that programmatically simulate and orchestrate user interactions with graphical user interfaces, including navigation, click and text-input sequences, and other GUI primitives; and analyzes interaction timing, element identification, and state synchronization to ensure reliable automated GUI behavior.
Automated human-computer interaction via GUI agents remains challenging due to fragmented evaluation criteria, heterogeneous architectures, and insufficiently characterized capabilities of large-model-driven agents. Method: We present the first unified capability framework for GUI agents, encompassing multimodal perception (OCR/VLM), neuro-symbolic reasoning, hierarchical task planning, and end-to-end reinforcement fine-tuning. We systematically classify and critically evaluate 15+ benchmarks, 30+ representative works, and eight architectural paradigms. Contribution/Results: Our work establishes a comprehensive technical landscape, introducing a reusable capability benchmarking paradigm, standardized evaluation protocols, and a forward-looking roadmap. We explicitly identify six open challenges—spanning robustness, generalization, compositional reasoning, efficiency, explainability, and real-world deployment—and outline concrete future research directions. This synthesis bridges theoretical foundations with practical engineering insights, advancing the systematic development of intelligent GUI agents.
To address the high learning barrier and inefficient documentation lookup associated with command-line interfaces (CLIs), this paper introduces GUIde—the first AI-powered system that automatically transforms unstructured Unix manual (man) pages into interactive graphical user interfaces (GUIs). Methodologically, GUIde leverages natural language processing and large language models to semantically parse man pages, extract parameter schemas, constraints, and usage examples, and generate executable interface specifications; a lightweight frontend engine then renders these specifications into responsive, visual GUIs in real time. Its core contribution is the first end-to-end, fully automated mapping from non-structured man text to functionally complete GUIs—requiring no manual annotation or tool-specific adaptation. Evaluated on a real-world corpus of 52 common Unix commands, GUIde achieves 96.3% coverage of valid parameter combinations and improves user task completion rates by 41.7%, effectively bridging the interaction paradigm gap between CLI and GUI environments.
This work proposes a vision-based GUI automation system centered on explicit task planning to address the tendency of existing agents to deviate from user intent in dynamic interfaces and their lack of transparent, intervenable planning mechanisms. By treating task plans as persistent, inspectable, and editable external artifacts, the system adopts a planning-execution decoupled architecture that integrates multimodal visual inputs, screenshot-anchored interventions, and natural language guidance. This design enables real-time monitoring and localized correction of execution trajectories. Experimental results demonstrate that the approach effectively recovers from the majority of automation failures, substantially enhancing the system’s transparency, controllability, and adaptability in complex, evolving graphical user interfaces.
Web applications exhibit high interface dynamism and complex navigation and form interactions, posing significant challenges for end-to-end test case generation—particularly low coverage and poor robustness. To address this, we propose a navigation modeling approach that integrates screen transition graphs with large language models (LLMs) to enable high-precision path exploration. We further design a state-graph-based automated testing framework for conditional forms, supporting dynamic DOM parsing and interaction modeling. Additionally, we construct the first dedicated benchmark dataset for form-interaction testing. Experimental evaluation across diverse, complex web applications demonstrates substantial improvements in test coverage and path discovery accuracy—especially for branching navigation and conditional form-filling tasks. Our approach establishes a scalable, empirically evaluable paradigm for web application reliability testing.
This paper addresses the ambiguous human–AI collaboration relationships and insufficient exploration of interaction paradigms in hybrid active visual analytics systems. Adopting a qualitative approach integrating bibliometric analysis and thematic coding, it synthesizes two decades of literature to construct a novel three-dimensional classification framework—spanning collaborative objectives, automation levels, and human roles—and proposes the first consensus-based operational definition of hybrid active visual analytics. The study identifies critical limitations: conceptual inconsistency, narrow interaction patterns, and imbalanced research distribution across domains and time periods. It further provides the first systematic characterization of evolutionary trajectories of real-world practices across eras and application scenarios. The findings establish a theoretical foundation, design guidelines, and a developmental roadmap for advancing human–AI collaborative visual analytics, thereby filling a foundational gap in modeling collaboration paradigms within hybrid active systems.
Existing AutoML systems for novice users prioritize algorithmic sophistication over usability, trust, and interpretability, hindering effective adoption. Method: This paper proposes an end-to-end abstracted pipeline tailored for novices, spanning data ingestion, guided configuration, training, evaluation, and inference. Grounded in four human-centered design principles—ensuring first-model success to boost self-efficacy, providing explanations to foster accurate mental models, applying contextual abstraction to maintain the zone of proximal development, and enhancing predictability and safety to strengthen perceived control—we employed user-centered design to develop a prototype and conducted a controlled 24-participant user study. Contribution/Results: UEQ (User Experience Questionnaire) results confirmed that all participants successfully built models, with significantly positive ratings for usability, trust, and comprehensibility. Domain-expert evaluations further corroborated the effectiveness of the design principles in supporting novice skill development.
This work addresses the lack of fair comparison between graphical user interface (GUI) and command-line interface (CLI) agents due to confounding variables such as interaction modalities, task definitions, initial states, action spaces, and verification mechanisms. To isolate performance bottlenecks, the authors construct a matched benchmark comprising 440 desktop tasks spanning 18 applications and 12 workflow categories, enforcing identical task objectives, initial conditions, and final-state validation while restricting agents to their native action sets. Under this rigorously controlled setting, GUI agents are found to be limited by the reliability of long-horizon visual interactions, whereas CLI agents suffer from insufficient skill-interface coverage. Based on this insight, the paper introduces a verifier-guided skill augmentation method. Experiments show that the best GUI agent achieves a 59.1% full-task success rate, the baseline CLI agent 48.2%, and the augmented CLI agent improves to 69.3%, confirming skill coverage as the critical limiting factor.
Users often fail to recognize automation opportunities in routine web interactions, and existing large language model–based approaches are costly and require explicit task specification. This work proposes an unobtrusive automation framework that passively monitors browser activity to automatically identify programmable, repetitive interaction patterns and generates installable automation scripts, which users can review and refine using natural language. The approach introduces the first mechanism for discovering automation opportunities without active user involvement, integrating behavioral tracking, pattern recognition, program synthesis, and a large language model interface. A user study demonstrates that the system uncovers substantially more automation opportunities than users identify on their own; most discovered automations align with real-world workflows and are rated as useful, yielding high user willingness to adopt.