Beyond Task Completion: Measuring Interaction Cost in Terminal User Interfaces

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the difficulty of quantifying interaction costs in terminal user interfaces and the overreliance of traditional evaluations on task completion rates. To overcome these limitations, it proposes the Agent-as-a-User paradigm alongside the TUINaut toolchain. This approach leverages large language model agents to simulate users executing authentic tasks while recording interaction trajectories, introducing a reproducible measurement framework grounded in independent oracle verification that effectively disentangles interpretation from operational costs. Across experiments involving 84 real-world tasks, the method demonstrates significant correlations with human interaction costs, revealing latent interaction disparities and usability deficiencies in generative interfaces that high success rates obscure. Ultimately, this work establishes the proposed paradigm as a standard methodology for evaluating the usability of task-oriented interfaces.
📝 Abstract
Large language models (LLMs) are increasingly used through terminal user interfaces (TUIs), yet task completion alone does not capture how difficult an interface is to understand and operate. Existing human assessments and LLM-generated ratings or reports do not provide repeatable measurements of interaction effort grounded in verified task execution. We propose Agent-as-a-User, an evaluation paradigm that places an LLM agent in the user role. Agent-KLM separates agent-side interpretation and operation costs and relates them structurally to human interaction. TUINaut operationalizes the paradigm by recording interaction trajectories and verifying outcomes with actor-independent oracles. Using TUINaut-Bench, we study 84 tasks across 15 real-world TUIs with six human participants and five observation-evaluator configurations, and evaluate 45 TUIs generated by three leading stacks. Similar aggregate success rates mask differences in which tasks humans and agents complete, whereas interpretation and operation costs follow correlated task rankings, especially for operation. This pattern persists across configurations. Generated TUIs can implement correct functionality while still requiring usability improvements. These results position Agent-as-a-User as a repeatable, execution-grounded paradigm for measuring task-based usability.
Problem

Research questions and friction points this paper is trying to address.

Terminal User Interfaces
Interaction Cost
Usability Evaluation
Task Completion
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agent-as-a-User
Interaction Cost
Terminal User Interfaces
Agent-KLM
Usability Evaluation
🔎 Similar Papers
No similar papers found.