Score
Designs and implements operational definitions and measurement frameworks for “experience quality,” including the selection and construction of quantitative and qualitative metrics, instruments, and scoring rules. Builds analyses and validation procedures that combine behavioral, observational, and self-report data to quantify, threshold, and interpret the quality of an experience for evaluation or comparison.
In data trading, expert-dependent Data Quality Assessment (DQA) impedes cross-organizational consensus, while practitioner experience heterogeneity exacerbates assessment bias. To address this, we conduct the first empirical study integrating eye-tracking with controlled comparative experiments to quantify how domain expertise influences perception and interpretation of quality metadata. Building on these findings, we propose a hierarchical, experience-adaptive DQA support paradigm and develop an automated tool for generating interpretable, multidimensional data quality metadata. The tool embeds explainable AI principles and integrates syntactic, semantic, and contextual quality indicators. Experimental evaluation demonstrates statistically significant reductions in DQA misclassification rates (p < 0.01) and improved usability for novice users. This work provides both theoretical foundations and empirical evidence for building trustworthy, transparent, and broadly deployable data quality infrastructure.
This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.
Existing foundation model evaluation methods—including human-centered paradigms—overemphasize output correctness while neglecting the dynamic coupling between response quality and user perception during interaction, thus failing to uncover the underlying mechanisms of user experience (UX). Method: We propose QoNext—the first framework to adapt the Quality of Experience (QoE) paradigm from networking and multimedia domains to large language model evaluation. Through controlled human-perception experiments, we systematically collect fine-grained subjective ratings across diverse configurations, establishing the first QoE-oriented benchmark dataset; we then train a high-accuracy UX prediction model using measurable system parameters. Contribution/Results: QoNext transcends traditional static-output evaluation by introducing a quantitative UX assessment framework that jointly models interaction dynamics and response quality. It enables proactive, interpretable, and optimization-aware evaluation of foundation model QoE, providing empirically grounded, deployment-ready tuning guidance for model development and productization.
Existing performance measurement frameworks struggle to simultaneously satisfy customizability, interpretability, and mathematical tractability in interdisciplinary contexts. Method: This paper proposes a goal-oriented, customizable metric construction framework featuring a novel “base metric–auxiliary metric” dichotomy. Integrating utility theory and multi-criteria decision analysis, it introduces an uncertainty-aware utility function and establishes a systematic metric decomposition–synthesis workflow. Contributions: (1) It reduces reliance on complex mathematical formalisms, enhancing applicability under resource constraints or high uncertainty; (2) it ensures metric transparency, traceability, and domain adaptability; and (3) it enables quantitative assessment of goal attainment, real-time progress monitoring, and downstream statistical modeling and decision optimization. The framework has been empirically validated across diverse disciplines, demonstrating generality and extensibility.
Existing XAI evaluation methods over-rely on single-explanation assessments and lack user-centered, quantitative tools. Method: We propose the first psychometrically grounded scale for measuring XAI experience quality (XEQ), grounded in a four-dimensional theoretical framework—learnability, utility, satisfaction, and engagement—to overcome limitations of static, point-in-time evaluations. XEQ was developed following rigorous scale development protocols, including expert content validity assessment and large-scale empirical validation (N=1,247) to establish discriminant and structural validity. Results: XEQ demonstrates excellent reliability and validity (Cronbach’s α = 0.92; CFI = 0.96), enabling systematic, multi-turn, and personalized XAI interaction assessment. As the first standardized, multidimensional, and reproducible measurement instrument for human-centered XAI evaluation, XEQ advances explainable AI research from a technology-centric to a user experience–driven paradigm.
This study addresses the persistent challenge in software engineering research of empirically validating theories due to the absence of systematic, reproducible operationalization methods. To bridge this gap, the authors propose an integrated methodological framework that combines Sjøberg’s operationalization approach with Dubin’s theory-building methodology, offering the first evidence-driven and replicable guide for operationalizing theoretical constructs in software engineering. The approach systematically translates abstract theories into measurable forms by rigorously defining variables, selecting appropriate indicators, and deriving non-causal assumptions. The utility of the framework is demonstrated through its application to a theory on DevOps team classification. The resulting methodology provides researchers with a robust foundation for conducting verifiable theoretical studies while simultaneously offering practitioners actionable, theory-informed insights.
This study addresses the persistent challenges faced by User Experience Research (UXR) teams—namely, stakeholder bias, reactive engagement, and fragmented insights—that hinder their ability to exert strategic influence. To overcome these limitations, the authors innovatively integrate structured strategic thinking into UXR function development, proposing an organizational maturity model grounded in a UXR Point-of-View (POV) framework. Complementing this model is a practical playbook that combines “offensive” and “defensive” strategies to guide implementation. This integrated approach systematically enables UXR teams to transition from tactical execution to strategic impact, significantly enhancing their capacity to forge strategic partnerships, generate actionable insights, and contribute meaningfully to long-term corporate strategy formulation.
This study addresses the prevalent ambiguity, inconsistency, and incompleteness in articulating explainability requirements for AI systems due to a lack of standardized specifications. Through a structured literature review and interviews with developers, the authors identify a set of explainability quality attributes, which are then refined via a large-scale survey of practitioners into ten core attributes. For the first time, these attributes are translated into a prioritized, actionable guideline for writing explainability requirements. Building on this foundation, the authors design a lightweight, iterative requirements engineering workflow augmented by a large language model to assist in requirement generation. An accompanying web-based tool reduces average requirement drafting time by 23.5%, and user evaluations indicate that the generated requirements match or slightly exceed manually written ones in terms of implementability and textual quality.
Traditional large language model (LLM) benchmarks often fail to capture real-world user experience, leading practitioners to rely on informal “vibe checks” that lack systematicity and reproducibility. This work formalizes vibe checks into a two-stage evaluation framework: personalized prompt generation grounded in user preferences, followed by subjective perception assessment. The authors instantiate this approach in a prototype benchmark and validate its efficacy through user studies, social media data analysis, and programming task experiments. Results demonstrate that the proposed method significantly alters model rankings compared to conventional metrics, effectively bridging the gap between standardized evaluations and actual user experience. By integrating personalized inputs with subjective judgment, this study introduces a novel evaluation paradigm that better reflects how users interact with and perceive LLMs in practice.