user study methodology

Designing and executing user-facing experiments to measure usability, perceived value, timing-of-feedback effects, and task-specific metrics (e.g., temporal consistency vs. reactivity), including choosing measurements and evaluation protocols.

userstudymethodology

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

ReVISit 2: A Full Experiment Life Cycle User Study Framework

Aug 05, 2025
ZC
Zach Cutler
🏛️ University of Utah | WPI

Visualization user studies face challenges including tool fragmentation, poor reproducibility, and inadequate support for complex interaction design. This paper introduces VisExp—a lightweight, open-source, browser-based framework that supports the full experimental lifecycle: design, pilot testing, data collection, analysis, and dissemination. VisExp uniquely integrates technical capabilities with socio-technical support mechanisms, enabling fine-grained interaction logging and precise behavioral replay. It provides automated data acquisition, an integrated analysis toolkit, and collaborative development interfaces. Furthermore, VisExp fosters a sustainable community ecosystem through extensible architecture and shared best practices. The framework has been successfully deployed in multiple ACM SIGCHI and IEEE VIS conference papers and rigorously replicated across independent research teams. Empirical evaluation demonstrates significant improvements in experimental efficiency, analytical transparency, and methodological reproducibility.

Addresses reproducibility and nuanced design challenges in experimentsFacilitates full lifecycle of online visualization user studiesSupports design, execution, and analysis of browser-based studies

HCI scale development has long suffered from nonstandardized processes, poor construct-theory alignment, and low item reuse rates. This paper introduces the first interactive support system integrating large language models (LLMs) with a structured, empirically grounded measurement knowledge base, enabling a closed-loop workflow: construct identification → theory-informed custom definition → context-aware item generation. The system retrieves theoretically appropriate constructs from a literature-anchored database and leverages LLMs to generate semantically coherent, domain-specific items, supporting human-AI co-refinement. Its key innovation lies in the deep coupling of LLMs with an evidence-validated construct–item relational database, shifting scale development from experience-driven practice toward evidence-enhanced collaborative measurement. Experiments show a 62% reduction in design time, a 3.1× increase in item reuse, and significantly improved theoretical fidelity; expert evaluations across multiple rounds confirm ≥92% contextual appropriateness. The system has been integrated into a prototype HCI research workflow.

Improving rigor and efficiency in HCI measurement designLeveraging LLMs and prior literature for construct developmentStandardizing measurement item design process for researchers

This work addresses the lack of a systematic framework to guide experimental design decisions in replication studies. It proposes the first multidimensional design space framework specifically tailored for replication research, conceptualizing replication as a pairwise comparison problem. The framework structures replication planning and analysis through four practical dimensions—task, data, method, and metrics—and three comparative levels: micro, meso, and macro. By offering actionable design guidelines and a comprehensive taxonomy, it enables both prospective planning and retrospective evaluation of replication efforts. Empirical case studies in visualization and human-computer interaction demonstrate the framework’s effectiveness in enhancing the rigor of replication designs and improving the comparability of evaluation outcomes.

design spaceexperimental designHCI

This study addresses the degradation of GUI performance model validity in crowdsourced experiments caused by participants who disregard instructions or interact haphazardly. To mitigate this issue, the authors propose a pre-task screening mechanism based on a brief, main-task-like GUI interaction—such as image scaling and matching—administered prior to the primary task. Participants’ interaction errors are captured as continuous data quality signals, enabling dynamic thresholding to filter out low-quality contributors. Empirical evaluations on both mouse-based and smartphone platforms demonstrate that this approach substantially reduces the prevalence of anomalous behavior and significantly improves the goodness-of-fit and predictive accuracy of GUI performance models. The method establishes a novel paradigm for ensuring reliable data quality in crowdsourced human-computer interaction research.

crowdsourcingdata qualityGUI experiments

Initiating and Replicating the Observations of Interactional Properties by User Studies Optimizing Applicative Prototypes

Jul 18, 2025
GR
Guillaume Rivière
🏛️ Univ. Bordeaux | ESTIA-Institute of Technology

HCI research suffers from numerous context-dependent, non-replicable empirical findings. To address this, we propose *Interaction Cycle Diffraction*—the first method to formalize and compare user interaction behavior across experimental conditions using *interactional properties* (e.g., feedback latency, action reversibility, or mode-switching cost) as fundamental analytical units, rather than interface morphology. This framework systematically enables identification, extraction, and validation of reproducible interactional properties across diverse prototypes, technologies, tasks, and user populations. Through iterative user studies and prototype refinement, we demonstrate its utility in continuously optimizing design workflows and accumulating reusable empirical knowledge. Our work establishes the first reproducibility framework for interactional properties in ubiquitous UIs, offering a novel paradigm for building a theoretical taxonomy and empirical foundation for an interaction science. (138 words)

Formalizing user interaction observations for replicationOptimizing applicative prototypes to improve user interactionsStudying interactional properties across diverse conditions

Latest Papers

What's happening recently
View more

This work addresses the lack of standardized protocols in large language model (LLM) agent systems, which hinders reproducibility, comparative analysis, and the monitoring of evaluator preference coupling and its temporal measurement decay. To resolve this, the paper introduces the Evaluator Preference Coupling (EPC) protocol—a four-stage isolation framework that standardizes executor and evaluator configurations, policy task design, TTRL update rules, and metric computation, establishing the first RFC-style measurement framework for LLM agents. The protocol incorporates versioned reference snapshots and a composite naming convention, integrating metrics such as gamma, Jensen–Shannon divergence (JSD), expected calibration error (ECE), and Brier score to structure outputs and API metadata. The authors release Reference Snapshot v1.0, encompassing eight evaluation conditions and 122 experimental replicates across mainstream models including GPT-4o, Qwen, and DeepSeek, alongside fully open-sourced protocols, snapshots, and implementation code.

Evaluator Bias PropagationEvaluator Preference CouplingLLM Agent Systems

This study addresses the lack of systematic support for creating, managing, and deploying stimulus materials in visualization experiments—a gap that often leads to invalid results or wasted resources. Through semi-structured interviews with 19 visualization researchers, the work systematically examines practices and challenges across the full lifecycle of stimulus materials, from exploration and selection to deployment and analysis, integrating perspectives from user research and human factors engineering. The findings reveal, for the first time, a heavy reliance on manual processes and significant scalability limitations as core pain points in current workflows. Building on these insights, the study identifies key opportunities for improvement, including automated generation and intelligent validation of stimuli, thereby laying the groundwork for future directions such as AI-assisted stimulus design.

experimental designresearch challengesstimuli deployment

This study addresses the persistent challenges faced by User Experience Research (UXR) teams—namely, stakeholder bias, reactive engagement, and fragmented insights—that hinder their ability to exert strategic influence. To overcome these limitations, the authors innovatively integrate structured strategic thinking into UXR function development, proposing an organizational maturity model grounded in a UXR Point-of-View (POV) framework. Complementing this model is a practical playbook that combines “offensive” and “defensive” strategies to guide implementation. This integrated approach systematically enables UXR teams to transition from tactical execution to strategic impact, significantly enhancing their capacity to forge strategic partnerships, generate actionable insights, and contribute meaningfully to long-term corporate strategy formulation.

institutional barriersresearch function maturitystakeholder bias

Existing agent evaluation methods rely on static benchmarks, which struggle to capture realistic failure modes in multi-step dynamic interactions and lack effective metrics for assessing interaction quality and coverage. To address these limitations, this work proposes VISTA—the first hybrid user simulation framework that integrates both UI and API interactions—and introduces a six-dimensional evaluation metric suite to comprehensively measure interaction realism, capability coverage, and effectiveness. Experiments in e-commerce shopping and educational customer service scenarios demonstrate that VISTA significantly enhances the realism and comprehensiveness of agent evaluations, generating assessment outcomes that are both more authentic and more broadly representative than those produced by current approaches.

failure modesinteractive agent evaluationrealistic user behaviors

This study addresses the limitations of existing AI evaluation methods, which often fail to align with real-world user needs, contextual nuances, and local policies, while manual assessment remains difficult to scale. To bridge this gap, the authors propose an auditable and iterative, context-aware evaluation framework that integrates persona-driven test case generation, domain-specific scoring rubrics, and a hybrid adjudication mechanism combining human reviewers and LLM-based judges. Automated scoring is activated only when sufficient agreement between LLM judgments and human annotations is achieved. A three-week pilot across four organizations involving 108 annotated question-answer pairs demonstrates that the approach effectively balances policy alignment with scalable automation, enabling reliable end-to-end evaluation of AI systems.

AI evaluationcontextual alignmenthuman-aligned scoring

Hot Scholars

CZ

Chengzhi Zhang

Nanjing University of Science and Technology
Text MiningNatural Language ProcessingScience of Science
HD

Huseyin Dogan

Professor of Human Computer Interaction
Human Computer InteractionHuman Centred AIAssistive TechDigital Health
MK

Marcos Kalinowski

Professor, Pontifical Catholic University of Rio de Janeiro (PUC-Rio)
Empirical Software EngineeringAI EngineeringAI4SEHuman Aspects in Software Engineering
ZL

Zhicong Lu

Assistant Professor, George Mason University
HCIsocial computinglive streamingcreativity support
JK

JaeWon Kim

University of Washington
Human-Computer InteractionSocial Computing