usability testing

Methods for evaluating user interfaces and experiences (including accessibility and neuroinclusive design) through user studies, field validation, and automated assessments to verify coverage of requirements and practical support for workflows.

usabilitytesting

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Analysis of User Experience Evaluation Methods for Deaf users: A Case Study on a mobile App

Jul 30, 2025
AE
A. E. Fuentes-Cortázar
🏛️ University of Veracruz

Existing UX evaluation methods predominantly target hearing users, overlooking the distinct needs of Deaf individuals—particularly regarding sign language use, visual information processing, and cultural identity—resulting in biased and incomplete assessments. Method: This study employs qualitative analysis, multiple case studies, and participatory evaluation with Deaf users to systematically examine the accessibility and applicability of mainstream UX methods in mobile applications, uncovering structural limitations in communicative adaptation. Contribution/Results: We propose a reconceptualized UX evaluation framework grounded in three principles: visual primacy, sign-language friendliness, and cultural responsiveness. Drawing on accessibility design guidelines, we develop targeted adaptation strategies. Empirical findings demonstrate that conventional methods—without modification—fail to accurately capture Deaf users’ interaction experiences and authentic needs; conversely, methodologically optimized approaches significantly enhance assessment validity and inclusivity.

Adapting methods to reflect Deaf users' communication needsEvaluating UX methods for Deaf users' accessibilityIdentifying limitations in traditional UX evaluation approaches

Are UX evaluation methods truly accessible

Aug 11, 2025
AE
Andrés Eduardo Fuentes-Cortázar
🏛️ Universidad Veracruzana

Current user experience (UX) evaluation methods inadequately address accessibility for Deaf users, largely overlooking linguistic and perceptual diversity. Method: Through critical literature review, empirical testing, and multimodal UX assessment techniques, this study evaluates the real-world applicability of existing “deaf-friendly” evaluation protocols. Contribution/Results: Findings reveal that mainstream methods rely excessively on auditory input and linear cognitive assumptions, neglecting sign language as a primary language, visual-dominant perception, and bilingual communication patterns—leading to distorted data collection and participatory barriers. The study introduces the novel “communicative accessibility” framework, exposing systemic deficiencies across perception, comprehension, and feedback dimensions. It argues for methodological reconstruction grounded in Deaf language acquisition principles and interactional norms. This work provides both theoretical foundations and practical guidelines for developing inclusive, psychometrically sound UX evaluation paradigms for Deaf users.

Adapt methods to meet Deaf users' communication needsAssess accessibility of UX methods for Deaf usersIdentify limitations in traditional UX evaluation tools

Towards User-Focused Cross-Domain Testing: Disentangling Accessibility, Usability, and Fairness

Jan 11, 2025
MD
Matheus de Morais Lecca
🏛️ University of Calgary

In software testing, the user-centered concepts of fairness, usability, and accessibility suffer from conceptual conflation, ambiguous boundaries, and fragmented practice. Method: We conducted a three-tier systematic literature review (SLR) synthesizing findings from 12 recent SLRs (2014–2024) to develop the first integrated, user-focused testing framework unifying these three dimensions. Through conceptual mapping, cross-domain comparison, and meta-analysis, we clarified their theoretical distinctions and elucidated their interdependent implementation mechanisms in AI systems. Contribution/Results: We propose novel cross-cutting evaluation dimensions and actionable, integrated testing guidelines. Our key contribution is the first domain-agnostic, synergistic testing model—moving beyond single-dimension paradigms—and an operational governance pathway with concrete implementation recommendations for real-world adoption.

Fairness TestingSoftware DevelopmentUsability and Accessibility Testing

This study addresses the limitations of existing AI evaluation methods, which often fail to align with real-world user needs, contextual nuances, and local policies, while manual assessment remains difficult to scale. To bridge this gap, the authors propose an auditable and iterative, context-aware evaluation framework that integrates persona-driven test case generation, domain-specific scoring rubrics, and a hybrid adjudication mechanism combining human reviewers and LLM-based judges. Automated scoring is activated only when sufficient agreement between LLM judgments and human annotations is achieved. A three-week pilot across four organizations involving 108 annotated question-answer pairs demonstrates that the approach effectively balances policy alignment with scalable automation, enabling reliable end-to-end evaluation of AI systems.

AI evaluationcontextual alignmenthuman-aligned scoring

Traditional GUI usability evaluation relies heavily on expert reviews and user testing, which are costly and inefficient, while existing computational agents struggle to accurately assess usability. This work proposes uxCUA—a machine learning–based computational user agent that, for the first time, integrates computable usability metrics with large-scale, labeled UI interaction data to enable end-to-end prediction of usability scores. By prioritizing interaction flows and simulating human-like operations, uxCUA generates fine-grained and credible usability critiques. Notably, it achieves higher evaluation accuracy than larger-scale models and demonstrates effectiveness on both synthetic and real-world GUI interfaces.

automated evaluationcomputer use agentsgraphical user interfaces

Latest Papers

What's happening recently
View more

This study addresses the inefficiency of traditional requirement engineering approaches that rely on manual annotation for extracting usability requirements from user reviews. To overcome this limitation, the authors propose a prompt engineering method leveraging large language models (LLMs) guided by Nielsen’s ten usability heuristics. They introduce, for the first time, a specialized prompt template tailored specifically for usability requirements and construct a dual-annotated dataset comprising 300 user reviews across multiple application categories, labeled both manually and by LLMs. Experimental results demonstrate that, with carefully designed prompts, LLMs achieve an F-score comparable to human annotators in identifying usability-related non-functional requirements, thereby confirming the feasibility and cost-effectiveness of the approach while underscoring the critical role of prompt design in model performance.

large language modelsnatural language understandingrequirements engineering

This work addresses the limited accessibility of technical bug reports generated by automated accessibility testing tools, which often hinder comprehension and timely remediation by non-technical stakeholders. To bridge this gap, the authors propose HEAR—a novel framework that integrates empathetic perspectives and legal compliance awareness into accessibility report generation. By leveraging large language models, HEAR transforms raw accessibility logs into narrative reports through UI context reconstruction, injection of personas representing users with disabilities, and multi-layered reasoning. Evaluation on four real-world Android applications demonstrates that HEAR-generated reports preserve factual accuracy while significantly enhancing empathy, perceived urgency, persuasiveness, and awareness of legal risks—all without imposing additional cognitive burden on readers.

accessibility testingbug report generationempathetic communication

This work addresses the current lack of controllable benchmarks for evaluating the reliability and actionability of user experience (UX) critiques generated by large language models, particularly across diverse interface scenarios. The authors propose UXBench, the first benchmark that assesses critique quality through downstream repair efficacy. It comprises ten categories of locally executable web components and incorporates a guided browser exploration mechanism, requiring models to produce structured UX reports grounded in interaction evidence. Report quality is measured by whether downstream repair agents can effectively improve interfaces based on these reports. UXBench introduces interaction-evidence constraints, a multidimensional scoring scheme, and blind human validation. Experiments across eight state-of-the-art models reveal that UX evaluation capability remains significantly multidimensional and unsaturated, with notable disparities in actionability, repair effectiveness, component reliability, and adaptability across interface types.

actionabilitybenchmarklarge language models

This study addresses the lack of standardized randomized controlled trial (RCT) frameworks in artificial intelligence evaluation, which has led to inconsistent research designs and limited reproducibility and comparability of results. Integrating Shadish’s four validity framework with TOP transparency guidelines, and drawing on RCT methodologies from clinical medicine, economics, psychology, and software engineering, this work proposes a structured RCT framework for AI that uniquely places human performance at the core of evaluation. It incorporates causal inference, heterogeneity analysis, and assessments of practical significance, while explicitly addressing AI-specific challenges such as model versioning, human–AI interaction, contamination effects, and fairness. The framework articulates five guiding principles and 33 actionable recommendations, supported by a tiered transparency mechanism that substantially enhances the rigor, reproducibility, and cross-study comparability of AI evaluation research, thereby establishing a foundational methodology for the field.

AI evaluationcausal inferencehuman performance

Hot Scholars

HQ

Huamin Qu

Chair Professor, Hong Kong University of Science and Technology
Data visualizationHuman-Computer InteractionExplainable AIE-Learning
MC

Mark Colley

University College London
Automated DrivingAugmented RealityDriver-Vehicle InteractionAccessibility
JN

Jan-Niklas Voigt-Antons

Professor of Computer Science, University of Applied Science Hamm-Lippstadt
eXtended Reality (XR)immersive MediaUser Experience
DW

Dakuo Wang

Northeastern University
Human-AI CollaborationHuman-Centered AIHuman-Computer InteractionAI for Healthcare
PH

Pan Hui

Chair Professor, Nokia Chair in Data Science, FREng & IEEE Fellow (HKUST & University of Helsinki)
Ubiquitous ComputingMobile ComputingAugmented RealityData Science