🤖 AI Summary
Existing GUI task difficulty metrics predominantly rely on motor-based indicators (e.g., step count), neglecting cognitive load. Method: We propose the “Cognitive Chain” framework, decomposing pre-execution cognition into three stages—discovery, decision-making, and computation—and quantify the difficulty of each stage using information-theoretic measures. Leveraging large language models, we automatically extract cognitive chains from user interaction traces to construct a computable Cognitive Difficulty Index (CDI). Contribution/Results: This work establishes the first cognitively grounded model for GUI task difficulty, moving beyond behaviorist paradigms. Empirical evaluation demonstrates that CDI significantly predicts per-step completion time (R² = 0.46) and reveals substantial performance degradation in state-of-the-art GUI agents under high cognitive load—validating both theoretical rigor and practical utility for HCI and AI-agent evaluation.
📝 Abstract
Measuring GUI task difficulty is crucial for user behavior analysis and agent capability evaluation. Yet, existing benchmarks typically quantify difficulty based on motor actions (e.g., step counts), overlooking the cognitive demands underlying task completion. In this work, we propose Cognitive Chain, a novel framework that models task difficulty from a cognitive perspective. A cognitive chain decomposes the cognitive processes preceding a motor action into a sequence of cognitive steps (e.g., finding, deciding, computing), each with a difficulty index grounded in information theories. We develop an LLM-based method to automatically extract cognitive chains from task execution traces. Validation with linear regression shows that our estimated cognitive difficulty correlates well with user completion time (step-level R-square=0.46 after annotation). Assessment of state-of-the-art GUI agents shows reduced success on cognitively demanding tasks, revealing capability gaps and Human-AI consistency patterns. We conclude by discussing potential applications in agent training, capability assessment, and human-agent delegation optimization.