Score
Designing and executing user-facing experiments to measure usability, perceived value, timing-of-feedback effects, and task-specific metrics (e.g., temporal consistency vs. reactivity), including choosing measurements and evaluation protocols.
Visualization user studies face challenges including tool fragmentation, poor reproducibility, and inadequate support for complex interaction design. This paper introduces VisExp—a lightweight, open-source, browser-based framework that supports the full experimental lifecycle: design, pilot testing, data collection, analysis, and dissemination. VisExp uniquely integrates technical capabilities with socio-technical support mechanisms, enabling fine-grained interaction logging and precise behavioral replay. It provides automated data acquisition, an integrated analysis toolkit, and collaborative development interfaces. Furthermore, VisExp fosters a sustainable community ecosystem through extensible architecture and shared best practices. The framework has been successfully deployed in multiple ACM SIGCHI and IEEE VIS conference papers and rigorously replicated across independent research teams. Empirical evaluation demonstrates significant improvements in experimental efficiency, analytical transparency, and methodological reproducibility.
HCI scale development has long suffered from nonstandardized processes, poor construct-theory alignment, and low item reuse rates. This paper introduces the first interactive support system integrating large language models (LLMs) with a structured, empirically grounded measurement knowledge base, enabling a closed-loop workflow: construct identification → theory-informed custom definition → context-aware item generation. The system retrieves theoretically appropriate constructs from a literature-anchored database and leverages LLMs to generate semantically coherent, domain-specific items, supporting human-AI co-refinement. Its key innovation lies in the deep coupling of LLMs with an evidence-validated construct–item relational database, shifting scale development from experience-driven practice toward evidence-enhanced collaborative measurement. Experiments show a 62% reduction in design time, a 3.1× increase in item reuse, and significantly improved theoretical fidelity; expert evaluations across multiple rounds confirm ≥92% contextual appropriateness. The system has been integrated into a prototype HCI research workflow.
This work addresses the lack of a systematic framework to guide experimental design decisions in replication studies. It proposes the first multidimensional design space framework specifically tailored for replication research, conceptualizing replication as a pairwise comparison problem. The framework structures replication planning and analysis through four practical dimensions—task, data, method, and metrics—and three comparative levels: micro, meso, and macro. By offering actionable design guidelines and a comprehensive taxonomy, it enables both prospective planning and retrospective evaluation of replication efforts. Empirical case studies in visualization and human-computer interaction demonstrate the framework’s effectiveness in enhancing the rigor of replication designs and improving the comparability of evaluation outcomes.
This study addresses the degradation of GUI performance model validity in crowdsourced experiments caused by participants who disregard instructions or interact haphazardly. To mitigate this issue, the authors propose a pre-task screening mechanism based on a brief, main-task-like GUI interaction—such as image scaling and matching—administered prior to the primary task. Participants’ interaction errors are captured as continuous data quality signals, enabling dynamic thresholding to filter out low-quality contributors. Empirical evaluations on both mouse-based and smartphone platforms demonstrate that this approach substantially reduces the prevalence of anomalous behavior and significantly improves the goodness-of-fit and predictive accuracy of GUI performance models. The method establishes a novel paradigm for ensuring reliable data quality in crowdsourced human-computer interaction research.
HCI research suffers from numerous context-dependent, non-replicable empirical findings. To address this, we propose *Interaction Cycle Diffraction*—the first method to formalize and compare user interaction behavior across experimental conditions using *interactional properties* (e.g., feedback latency, action reversibility, or mode-switching cost) as fundamental analytical units, rather than interface morphology. This framework systematically enables identification, extraction, and validation of reproducible interactional properties across diverse prototypes, technologies, tasks, and user populations. Through iterative user studies and prototype refinement, we demonstrate its utility in continuously optimizing design workflows and accumulating reusable empirical knowledge. Our work establishes the first reproducibility framework for interactional properties in ubiquitous UIs, offering a novel paradigm for building a theoretical taxonomy and empirical foundation for an interaction science. (138 words)
This work addresses the lack of standardized protocols in large language model (LLM) agent systems, which hinders reproducibility, comparative analysis, and the monitoring of evaluator preference coupling and its temporal measurement decay. To resolve this, the paper introduces the Evaluator Preference Coupling (EPC) protocol—a four-stage isolation framework that standardizes executor and evaluator configurations, policy task design, TTRL update rules, and metric computation, establishing the first RFC-style measurement framework for LLM agents. The protocol incorporates versioned reference snapshots and a composite naming convention, integrating metrics such as gamma, Jensen–Shannon divergence (JSD), expected calibration error (ECE), and Brier score to structure outputs and API metadata. The authors release Reference Snapshot v1.0, encompassing eight evaluation conditions and 122 experimental replicates across mainstream models including GPT-4o, Qwen, and DeepSeek, alongside fully open-sourced protocols, snapshots, and implementation code.
This study addresses the lack of systematic support for creating, managing, and deploying stimulus materials in visualization experiments—a gap that often leads to invalid results or wasted resources. Through semi-structured interviews with 19 visualization researchers, the work systematically examines practices and challenges across the full lifecycle of stimulus materials, from exploration and selection to deployment and analysis, integrating perspectives from user research and human factors engineering. The findings reveal, for the first time, a heavy reliance on manual processes and significant scalability limitations as core pain points in current workflows. Building on these insights, the study identifies key opportunities for improvement, including automated generation and intelligent validation of stimuli, thereby laying the groundwork for future directions such as AI-assisted stimulus design.
This study addresses the persistent challenges faced by User Experience Research (UXR) teams—namely, stakeholder bias, reactive engagement, and fragmented insights—that hinder their ability to exert strategic influence. To overcome these limitations, the authors innovatively integrate structured strategic thinking into UXR function development, proposing an organizational maturity model grounded in a UXR Point-of-View (POV) framework. Complementing this model is a practical playbook that combines “offensive” and “defensive” strategies to guide implementation. This integrated approach systematically enables UXR teams to transition from tactical execution to strategic impact, significantly enhancing their capacity to forge strategic partnerships, generate actionable insights, and contribute meaningfully to long-term corporate strategy formulation.
Existing agent evaluation methods rely on static benchmarks, which struggle to capture realistic failure modes in multi-step dynamic interactions and lack effective metrics for assessing interaction quality and coverage. To address these limitations, this work proposes VISTA—the first hybrid user simulation framework that integrates both UI and API interactions—and introduces a six-dimensional evaluation metric suite to comprehensively measure interaction realism, capability coverage, and effectiveness. Experiments in e-commerce shopping and educational customer service scenarios demonstrate that VISTA significantly enhances the realism and comprehensiveness of agent evaluations, generating assessment outcomes that are both more authentic and more broadly representative than those produced by current approaches.
This study addresses the limitations of existing AI evaluation methods, which often fail to align with real-world user needs, contextual nuances, and local policies, while manual assessment remains difficult to scale. To bridge this gap, the authors propose an auditable and iterative, context-aware evaluation framework that integrates persona-driven test case generation, domain-specific scoring rubrics, and a hybrid adjudication mechanism combining human reviewers and LLM-based judges. Automated scoring is activated only when sufficient agreement between LLM judgments and human annotations is achieved. A three-week pilot across four organizations involving 108 annotated question-answer pairs demonstrates that the approach effectively balances policy alignment with scalable automation, enabling reliable end-to-end evaluation of AI systems.