Score
Designs, builds, or analyzes products, systems, interfaces, or workflows to ensure they enable intended users to complete tasks efficiently, effectively, and with satisfaction. This includes planning and conducting usability studies, heuristic evaluations, task and cognitive analyses, defining usability metrics, and iterating designs based on observed user behavior and feedback.
This study investigates whether large language models (LLMs) can bridge the gap between UX experts and non-experts in authoring user scenarios. In a controlled experiment, both groups authored scenarios with LLM assistance; outputs were evaluated via mixed methods—structured scoring and qualitative coding—assessing structural completeness, expressive clarity, and audience orientation. Results demonstrate, for the first time empirically, that LLMs significantly enhance non-experts’ performance: their scenarios achieve structural and clarity levels comparable to experts’, and—remarkably—surpass experts in articulating user perspectives. The findings validate LLMs as effective, democratized tools for requirements analysis and reveal their unique capacity to augment empathic user-centered expression. This work advances accessible UX practice by lowering barriers to rigorous scenario-based design.
This study investigates how the format of requirements representation—UML sequence diagrams versus plain text—affects the accuracy of requirements inspection, and examines the moderating roles of working memory capacity and mental rotation ability. Employing a crossover experimental design, the research integrates cognitive ability assessments with task performance data, analyzed via linear mixed-effects models and generalized linear models. Results reveal a significant three-way interaction among representation type and the two cognitive abilities: participants with higher cognitive capacities provided more accurate justifications for identified defects when supported by UML diagrams, yet paradoxically exhibited lower defect detection rates. These findings challenge the assumption that UML universally enhances inspection effectiveness and demonstrate that the benefits of multimodal representations are highly contingent on individual cognitive characteristics.
This study addresses the systemic support of macrocognitive functions—namely, event detection, sensemaking, adaptability, perspective shifting, and coordination—in human–AI teaming, moving beyond traditional usability-centered design paradigms. Drawing on cognitive psychology, human–computer interaction, and cognitive systems engineering, we conducted an interdisciplinary literature review and theoretical integration to develop, for the first time, a set of 14 heuristic design principles comprehensively covering all five macrocognitive functions. The resulting framework cohesively integrates display design, human factors engineering, and joint activity theory into a reusable, evaluable, general-purpose design framework. Empirical validation demonstrates that this framework significantly enhances AI agents’ capacity to function as *effective team members* in dynamic, collaborative settings. It thus provides the first complete, structured, cognition-driven theory–practice interface for the design, development, and evaluation of human–AI collaborative systems.
This study addresses the current lack of a systematic understanding of user interaction mechanisms with large language model–driven computer-use agents and the key design factors influencing their user experience (UX). Through a two-stage approach, the authors construct and empirically validate a UX design space for such agents. First, they synthesize findings from a literature review and expert interviews to develop a taxonomy encompassing dimensions such as user prompting, explainability, and user control. Second, they conduct a Wizard-of-Oz experiment across normal, error, and high-risk scenarios to observe user behaviors, revealing interdependencies among design dimensions and the diversity of user needs. This work presents the first systematically formulated and empirically validated UX design framework for LLM-driven agents, offering developers a structured and actionable foundation for design decisions.
Human-centered Requirements Engineering (HC-RE) integrates user cognition, emotions, and social interactions into the RE process through contributions from disciplines such as psychology, cognitive science, design thinking, and human-computer interaction. Despite growing interest, how these multidisciplinary contributions are structured and why they remain fragmented across the RE lifecycle is not well understood. This systematic mapping study analyzes 56 primary studies across seven dimensions, including RE phases, user involvement techniques, contributing disciplines, and evaluation methods. Results show that 70\% of approaches involve multidisciplinary contributions, yet only 39% have been empirically evaluated and 48% address only the elicitation phase. A cross-study analysis reveals a structural separation between two parallel integration traditions: a Cognitive-Formal (C-F) pathway grounded in goal-based frameworks and formal modeling, and a Participatory-Iterative (P-I) pathway grounded in scenario-based frameworks and iterative design. Each pathway has developed complementary strengths, but their near-total disconnection explains the persistent lifecycle concentration and theory-practice gap observed in the corpus. The findings identify the absence of translation mechanisms between human-centered artifacts and formal RE specifications as the field's primary structural gap, provide a structured research agenda organized into four priority tiers, and establish the empirical foundation for Experience-Centered Requirements Engineering, a direction in which user experience is explicitly operationalized as a first-class concern in requirements specification.
This study addresses the limitation that user interfaces generated by large language models typically satisfy functional requirements while lacking constraints derived from high-quality design principles. To overcome this, we construct a shared human-machine cognitive space and pioneer the translation of classical HCI design knowledge into version-controlled, executable declarative skill files. By integrating structured prompt engineering with software skill modularization techniques, these design principles are encoded as machine-readable skills and injected directly into the generation process. Consequently, this work enables the on-demand generation of specification-compliant user interfaces, advancing design compliance from static checklists toward a dynamic, executable, and open automated paradigm.
This study addresses the lack of domain-specific usability evaluation guidelines and systematic analytical tools for configurator user interfaces, which has led to inefficient and insufficiently expert assessments. To bridge this gap, the work proposes a semi-automated evaluation framework that leverages multimodal large language models (MLLMs) to jointly reason over visual and textual inputs. Grounded in 18 configurator-specific usability criteria, the framework automatically identifies usability issues, rates their severity, and generates actionable improvement suggestions. Experimental validation on 16 real-world configurators demonstrates that the approach reliably and efficiently diagnoses configurator-unique usability problems, substantially reducing manual effort while exhibiting strong scalability and adaptability across domains.
This study addresses the inefficiency of traditional requirement engineering approaches that rely on manual annotation for extracting usability requirements from user reviews. To overcome this limitation, the authors propose a prompt engineering method leveraging large language models (LLMs) guided by Nielsen’s ten usability heuristics. They introduce, for the first time, a specialized prompt template tailored specifically for usability requirements and construct a dual-annotated dataset comprising 300 user reviews across multiple application categories, labeled both manually and by LLMs. Experimental results demonstrate that, with carefully designed prompts, LLMs achieve an F-score comparable to human annotators in identifying usability-related non-functional requirements, thereby confirming the feasibility and cost-effectiveness of the approach while underscoring the critical role of prompt design in model performance.
This study addresses the high cost and expert dependency of traditional usability evaluations, which often hinder adoption by small development teams. It presents the first application of multimodal large language models (MLLMs) to automated usability analysis, leveraging screenshots and user interaction recordings to automatically detect usability issues based on Nielsen’s heuristics. The approach generates actionable improvement suggestions and incorporates severity-based prioritization to alleviate developers’ burden in determining issue importance. Findings from a user study indicate that software engineers perceive the high-priority recommendations as both high-quality and highly practical, demonstrating the method’s effectiveness as a low-cost, accessible complement to conventional usability assessment practices.
Current intelligent agents exhibit limited generalization capabilities when confronted with unseen or dynamically changing user interfaces, often failing to complete tasks robustly. This work presents the first systematic analysis of how Nielsen’s Ten Usability Heuristics impact agent performance, deriving interface design principles tailored specifically for agents and proposing safety-aware interface augmentation strategies. Through a controlled experimental environment—UI-Verse—combined with agent interaction modeling and dual human-agent evaluations, the study validates the efficacy of the proposed approach. Experimental results demonstrate that heuristic-based enhancements significantly improve both task success rates and efficiency for agents, with combined strategies yielding the best outcomes. Crucially, user studies confirm that these modifications do not compromise usability for human users.