Score
Designs and implements metrics and composite indicators that quantify institutional quality, including operationalizing subgroup moderator variables and computing comparative quality scores. Builds validation and statistical-testing workflows to assess measurement invariance across eras or contexts and to test moderation effects.
Existing research frequently suffers from model misspecification of formative constructs, and the absence of a consensus-based validation methodology leads scholars to erroneously apply reflective measurement frameworks, thereby compromising construct validity. Method: This paper introduces the first dedicated, multi-stage validation framework for formative constructs, integrating systematic literature review, descriptive statistics, multicollinearity diagnostics, and formative-model-specific tests to rigorously distinguish formative (causal) from reflective (effect) measurement logic. Contribution/Results: The framework ensures both theoretical rigor and practical feasibility, substantially enhancing the psychometric soundness and statistical integrity of formative indicators. It provides a reproducible, defensible methodological pathway for scale development and construct validation, directly addressing longstanding measurement challenges in behavioral and social science research.
Existing performance measurement frameworks struggle to simultaneously satisfy customizability, interpretability, and mathematical tractability in interdisciplinary contexts. Method: This paper proposes a goal-oriented, customizable metric construction framework featuring a novel “base metric–auxiliary metric” dichotomy. Integrating utility theory and multi-criteria decision analysis, it introduces an uncertainty-aware utility function and establishes a systematic metric decomposition–synthesis workflow. Contributions: (1) It reduces reliance on complex mathematical formalisms, enhancing applicability under resource constraints or high uncertainty; (2) it ensures metric transparency, traceability, and domain adaptability; and (3) it enables quantitative assessment of goal attainment, real-time progress monitoring, and downstream statistical modeling and decision optimization. The framework has been empirically validated across diverse disciplines, demonstrating generality and extensibility.
This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.
This study investigates the interrelationships among system quality (SQ), information quality (IQ), and service quality (SerQ) in information systems (IS), and their joint impact on customer satisfaction and operational efficiency. Method: Drawing on cross-industry survey data and structural equation modeling (SEM), the research empirically tests hypothesized causal pathways. Contribution/Results: The study first empirically validates a sequential mediation path: SQ → IQ → SerQ → user satisfaction. It further identifies SerQ as the strongest predictor of IS performance (Cronbach’s α = 0.953; KMO = 0.965), establishing it as a pivotal, integrative metric of overall IS effectiveness. Based on these findings, the study proposes an actionable IS quality evaluation framework centered on SerQ, offering both theoretical grounding and practical guidance for IS design, assessment, and continuous improvement.
This study addresses the frequent neglect of local cultural perspectives in existing automated evaluations of AI-generated images, particularly regarding “cultural appropriateness.” It introduces a novel evaluation framework that deeply integrates diverse community participation from the outset, collaborating with blind and visually impaired individuals in the UK and residents of Kerala and Tamil Nadu in India to systematically translate lived cultural experiences and community concerns into actionable assessment dimensions. Leveraging multimodal large language models as judges (LLM-as-a-judge), the approach operationalizes community consensus into structured scoring rules, enabling automated evaluation of cultural appropriateness. The work not only establishes a conceptual framework grounded in community values and demonstrates its feasibility but also exposes critical limitations in current AI models’ understanding of cultural context.
This study addresses the lack of systematic evaluation of data quality tools with respect to their measurement capabilities and integration with large language models (LLMs). It presents the first multidimensional assessment framework grounded in real-world enterprise use cases, systematically evaluating six prominent tools—including open-source solutions such as Great Expectations and Deequ, as well as commercial platforms like Informatica and Experian—across dimensions including rule definition, duplicate detection, metric aggregation, and uncertainty handling, along with their LLM integration mechanisms. The findings reveal that commercial tools offer more comprehensive functionality and初步 support for LLM-assisted rule generation, whereas open-source tools provide greater flexibility at the cost of higher implementation effort. Notably, none of the evaluated tools currently enable direct LLM-based data validation. This work provides empirical guidance for selecting data quality tools and advancing their integration with LLMs.
This study addresses the current lack of systematic approaches for evaluating artificial intelligence’s adaptability and comprehension across diverse cultural contexts. Drawing on measurement theory, it introduces—for the first time—the validity framework from psychometrics into the assessment of AI cultural competence, thereby disentangling the construct of “cultural intelligence” from its operationalization. The work proposes a modular and extensible evaluation paradigm that integrates cultural dimension modeling, indicator design, data collection, and assessment protocols. By delineating core competency domains and their corresponding measurable indicators, this research establishes a theoretical and methodological foundation for large-scale, systematic evaluation of AI systems’ cultural adaptability.
This study addresses the ambiguity and inconsistency in evaluation criteria for software engineering replication studies, which have led to contradictory interpretations and uncertainty in reported results. Through a systematic review of ten replication studies published between 2021 and 2025, combined with qualitative content analysis, statistical principles, and modeling of measurement uncertainty, this work is the first to uncover the heterogeneity and lack of standardized practices in current evaluation approaches. Building on these insights, the paper proposes a unified evaluation framework that integrates statistical theory, methodological rigor, and measurement theory. Empirical illustration demonstrates that the framework effectively enhances the transparency, consistency, comparability, and reliability of replication studies in software engineering.