Score
Design and document precise operational definitions that translate theoretical constructs into measurable variables, indicators, and proxy metrics; this includes specifying variable coding schemes, measurable inclusion/exclusion criteria, and mapping indicators to available data. Build annotation and measurement protocols, guidelines, and validation procedures to generate, label, and empirically evaluate the resulting metrics and documentation.
This study addresses the persistent challenge in software engineering research of empirically validating theories due to the absence of systematic, reproducible operationalization methods. To bridge this gap, the authors propose an integrated methodological framework that combines Sjøberg’s operationalization approach with Dubin’s theory-building methodology, offering the first evidence-driven and replicable guide for operationalizing theoretical constructs in software engineering. The approach systematically translates abstract theories into measurable forms by rigorously defining variables, selecting appropriate indicators, and deriving non-causal assumptions. The utility of the framework is demonstrated through its application to a theory on DevOps team classification. The resulting methodology provides researchers with a robust foundation for conducting verifiable theoretical studies while simultaneously offering practitioners actionable, theory-informed insights.
This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.
Data scientists frequently lack systematic guidance when operationalizing ambiguous concepts (e.g., “writing authenticity,” “medical need”) into model-ready proxy target variables. To address this, we conducted semi-structured interviews with 15 data scientists across education and healthcare domains, followed by cross-domain thematic coding. We propose the “assemblage metrics” framework, identifying five core design criteria: validity, simplicity, predictiveness, portability, and resource efficiency. Our analysis reveals an iterative, problem-reconstruction–driven practice in which target variables are dynamically negotiated through trade-offs among these criteria. This work offers the first systematic characterization of such trade-offs in proxy target construction. It contributes a theoretically grounded framework and methodological tools for HCI, CSCW, and machine learning communities to support principled, transparent, and trustworthy predictive modeling—bridging conceptual abstraction with operationalizable measurement.
本文通过探索性因子分析(EFA)和验证性因子分析(CFA)评估面向对象、类级代码质量度量的构建有效性,确定了24个有效的代码质量度量。
Existing performance measurement frameworks struggle to simultaneously satisfy customizability, interpretability, and mathematical tractability in interdisciplinary contexts. Method: This paper proposes a goal-oriented, customizable metric construction framework featuring a novel “base metric–auxiliary metric” dichotomy. Integrating utility theory and multi-criteria decision analysis, it introduces an uncertainty-aware utility function and establishes a systematic metric decomposition–synthesis workflow. Contributions: (1) It reduces reliance on complex mathematical formalisms, enhancing applicability under resource constraints or high uncertainty; (2) it ensures metric transparency, traceability, and domain adaptability; and (3) it enables quantitative assessment of goal attainment, real-time progress monitoring, and downstream statistical modeling and decision optimization. The framework has been empirically validated across diverse disciplines, demonstrating generality and extensibility.
This study addresses the high cost and low efficiency of current ontology extension practices, which heavily rely on manual effort due to the underutilization of domain knowledge implicitly embedded in operational metrics. To overcome this limitation, the work proposes the first context-aware ontology extension framework that systematically leverages structured operational metrics as a source of contextual information. The framework formulates ontology extension as three subtasks: parent class prediction, relationship type prediction, and data property assignment, and integrates natural language processing with knowledge graph techniques to generate automated suggestions. Experimental evaluation on four cybersecurity ontologies demonstrates that the proposed approach significantly outperforms baseline methods relying solely on ontology-internal context, particularly in relationship type prediction and data property assignment, thereby effectively reducing the cost of ontology maintenance.
研究解决数据仓库中缺乏变量级元数据的问题,通过提出一种与DDI-CDI模型兼容的元数据应用配置文件方法来丰富元数据并进行模型-数据一致性检查。
This study addresses the growing challenge posed by the widespread involvement of AI agents in software development, which undermines the long-standing assumption that development artifacts are exclusively produced by human professionals—an assumption underpinning traditional software metrics. The work systematically exposes how AI-generated traces compromise the foundational premises of established software measurement practices, thereby threatening the validity of prior empirical conclusions. To confront this issue, the authors propose an AI-augmented, systematic replication methodology that integrates modern data analytics with empirical software engineering techniques to rigorously re-evaluate key findings. The project advances a dynamic, reproducible, and sustainable measurement paradigm capable of adapting to evolving data ecosystems, offering a robust and timely framework for software metrics in the AI era.
One aspired outcome of empirical research on quantitative data is a variance theory, i.e., a quantification of the effect of an independent on a dependent variables. The validity of variance theories stems from the synthesis of multiple pieces of evidence, which increases its validity beyond the findings of a single study. However, research synthesis in SE is rare and if done mostly limited to purely narrative syntheses. At best, researchers perform meta-analyses to synthesize variance theories from several quantitative results. But even meta-analyses only produce reliable results when synthesizing exact replications yet fail to generalize from variations. We aim to extend the frontier of research synthesis beyond the state-of-the-art to systematically manage empirical evidence and its evolution. We apply method engineering to construct a framework for research synthesis from proven, individual method fragments. The framework allows researchers to put new evidence in a clear relation to an existing body of evidence and systematically expand knowledge about a studied phenomenon. We demonstrate the application of this framework to two fields of research by explicitly modeling the relationship between existing pieces of evidence. The framework puts three types of evolution of evidence into relation: (1) replications investigate the same hypothesis in a new context to improve external validity, (2) revisions challenge an existing hypothesis to improve internal validity, and (3) reanalyses replace analysis methods to improve conclusion validity. Through a systematic evolution of evidence and clear assessment criteria for each dimension of validity, the proposed framework can determine the frontier of a field of research. The framework provides a perspective to systematically evolve empirical evidence in SE, supporting more constructive and productive advances in our field.
This study addresses the lack of implementation guidelines for ISO data quality standards, which hinders their practical adoption. We systematically categorize ISO metrics and formulate executable specifications, developing dqmeasure, an open-source Python library for automated assessment. The core innovation lies in a novel method that automatically learns parameters from reference data, eliminating reliance on manual rules. Experimental results demonstrate that the proposed metrics decrease monotonically as data errors increase and exhibit strong correlation with downstream machine learning performance. Furthermore, the approach supports linearly scalable monitoring. This work provides an effective tool for the automated evaluation of data quality.