Score
Designs and conducts empirical studies and hands-on tests to evaluate and improve human factors, usability, and user acceptance of interactive systems — including websites, APIs, SDKs, tools, and deployment sites. Creates test protocols and acceptance criteria, selects and applies usability metrics and methods, collects and analyzes user performance and satisfaction data (errors, task completion, learnability, trust), and iterates product designs based on observed workflows and feedback.
Traditional GUI usability evaluation relies heavily on expert reviews and user testing, which are costly and inefficient, while existing computational agents struggle to accurately assess usability. This work proposes uxCUA—a machine learning–based computational user agent that, for the first time, integrates computable usability metrics with large-scale, labeled UI interaction data to enable end-to-end prediction of usability scores. By prioritizing interaction flows and simulating human-like operations, uxCUA generates fine-grained and credible usability critiques. Notably, it achieves higher evaluation accuracy than larger-scale models and demonstrates effectiveness on both synthetic and real-world GUI interfaces.
This work addresses the challenge of transforming unstructured feedback from simulated user agents in usability testing into actionable user experience insights. To this end, it introduces UXCascade, an interactive analysis tool that pioneers a multi-level analytical framework integrating user personas, task objectives, and usability issues to link agent reasoning traces with specific interface problems. The system supports exploratory analysis—from macro-level patterns to micro-level refinements—through structured overviews, reasoning trace visualization, annotation views, and interactive editing capabilities. A user study demonstrates that UXCascade effectively integrates into existing UX workflows, facilitating rapid iteration during early design stages and yielding high-value, actionable feedback.
This study investigates whether large language models (LLMs) can bridge the gap between UX experts and non-experts in authoring user scenarios. In a controlled experiment, both groups authored scenarios with LLM assistance; outputs were evaluated via mixed methods—structured scoring and qualitative coding—assessing structural completeness, expressive clarity, and audience orientation. Results demonstrate, for the first time empirically, that LLMs significantly enhance non-experts’ performance: their scenarios achieve structural and clarity levels comparable to experts’, and—remarkably—surpass experts in articulating user perspectives. The findings validate LLMs as effective, democratized tools for requirements analysis and reveal their unique capacity to augment empathic user-centered expression. This work advances accessible UX practice by lowering barriers to rigorous scenario-based design.
This work proposes an end-to-end agent framework that addresses the limitations of traditional web usability evaluation, which relies on time-consuming user studies and expert reviews ill-suited for rapid iterative development. The framework uniquely integrates multimodal GUI perception with simulated user behavior profiling to interact directly with live web pages without requiring DOM parsing. It incorporates structured usability protocols—including the System Usability Scale (SUS), Single Ease Question (SEQ), and think-aloud methods—to automatically generate standardized user experience reports. Built upon the Avenir-Web architecture, the approach leverages joint visual-semantic modeling and multimodal action grounding to significantly enhance the automation and scalability of usability testing, thereby empowering developers to efficiently create highly usable web interfaces.
HCI research suffers from numerous context-dependent, non-replicable empirical findings. To address this, we propose *Interaction Cycle Diffraction*—the first method to formalize and compare user interaction behavior across experimental conditions using *interactional properties* (e.g., feedback latency, action reversibility, or mode-switching cost) as fundamental analytical units, rather than interface morphology. This framework systematically enables identification, extraction, and validation of reproducible interactional properties across diverse prototypes, technologies, tasks, and user populations. Through iterative user studies and prototype refinement, we demonstrate its utility in continuously optimizing design workflows and accumulating reusable empirical knowledge. Our work establishes the first reproducibility framework for interactional properties in ubiquitous UIs, offering a novel paradigm for building a theoretical taxonomy and empirical foundation for an interaction science. (138 words)
This work addresses the current lack of controllable benchmarks for evaluating the reliability and actionability of user experience (UX) critiques generated by large language models, particularly across diverse interface scenarios. The authors propose UXBench, the first benchmark that assesses critique quality through downstream repair efficacy. It comprises ten categories of locally executable web components and incorporates a guided browser exploration mechanism, requiring models to produce structured UX reports grounded in interaction evidence. Report quality is measured by whether downstream repair agents can effectively improve interfaces based on these reports. UXBench introduces interaction-evidence constraints, a multidimensional scoring scheme, and blind human validation. Experiments across eight state-of-the-art models reveal that UX evaluation capability remains significantly multidimensional and unsaturated, with notable disparities in actionability, repair effectiveness, component reliability, and adaptability across interface types.
This study addresses the limitations of existing AI evaluation methods, which often fail to align with real-world user needs, contextual nuances, and local policies, while manual assessment remains difficult to scale. To bridge this gap, the authors propose an auditable and iterative, context-aware evaluation framework that integrates persona-driven test case generation, domain-specific scoring rubrics, and a hybrid adjudication mechanism combining human reviewers and LLM-based judges. Automated scoring is activated only when sufficient agreement between LLM judgments and human annotations is achieved. A three-week pilot across four organizations involving 108 annotated question-answer pairs demonstrates that the approach effectively balances policy alignment with scalable automation, enabling reliable end-to-end evaluation of AI systems.
This study addresses the high cost and expert dependency of traditional usability evaluations, which often hinder adoption by small development teams. It presents the first application of multimodal large language models (MLLMs) to automated usability analysis, leveraging screenshots and user interaction recordings to automatically detect usability issues based on Nielsen’s heuristics. The approach generates actionable improvement suggestions and incorporates severity-based prioritization to alleviate developers’ burden in determining issue importance. Findings from a user study indicate that software engineers perceive the high-priority recommendations as both high-quality and highly practical, demonstrating the method’s effectiveness as a low-cost, accessible complement to conventional usability assessment practices.
This work addresses the challenges faced by resource-constrained software startup teams with limited user experience (UX) expertise in efficiently creating and evaluating low-fidelity prototypes. To this end, we propose SoftBoard, a web-based multi-agent system that integrates large language model–driven intelligent agents into the prototyping workflow for the first time, enabling an end-to-end pipeline from requirement elicitation to automated generation of low-fidelity prototypes. The system incorporates an embedded evaluation mechanism based on usability heuristic rules and unifies prototype editing, team collaboration, and AI-assisted functionalities within a single platform, substantially reducing reliance on specialized UX knowledge. Preliminary feasibility studies demonstrate that SoftBoard effectively standardizes and streamlines the minimum viable product (MVP) development process.