metrics

Designs and implements quantitative measures and measurement frameworks to evaluate and monitor the performance, quality, and behavior of systems, models, or processes; defines metric formulas, aggregation rules, baselines, thresholds, and validation tests. Analyzes metric properties (e.g., bias, variance, sensitivity), instrumentation and data collection, computes and visualizes results, and interprets them to support comparison and decision-making.

metrics

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.99
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$199K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Root/Additional Metric (RoAM) framework: a guide for goal-centred metric construction

Jul 02, 2025
LE
Luke E. B. Goodyear
🏛️ Queen’s University Belfast

Existing performance measurement frameworks struggle to simultaneously satisfy customizability, interpretability, and mathematical tractability in interdisciplinary contexts. Method: This paper proposes a goal-oriented, customizable metric construction framework featuring a novel “base metric–auxiliary metric” dichotomy. Integrating utility theory and multi-criteria decision analysis, it introduces an uncertainty-aware utility function and establishes a systematic metric decomposition–synthesis workflow. Contributions: (1) It reduces reliance on complex mathematical formalisms, enhancing applicability under resource constraints or high uncertainty; (2) it ensures metric transparency, traceability, and domain adaptability; and (3) it enables quantitative assessment of goal attainment, real-time progress monitoring, and downstream statistical modeling and decision optimization. The framework has been empirically validated across diverse disciplines, demonstrating generality and extensibility.

Combines decision analysis and utility theory to quantify goal achievementDevelops a framework for constructing customizable performance metrics across disciplinesDivides criteria into root and additional groups for flexible metric design

Existing software engineering metrics often fail to effectively support critical development decisions—such as whether refactoring is necessary or whether testing is sufficient—thereby limiting their practical utility. This work addresses this gap by systematically introducing metrological principles (the science of measurement) into the domain of software measurement for the first time. It proposes a metrology-informed approach to metric modeling and evaluation, establishing a rigorous scientific foundation for the design of software metrics. By grounding metric development in established measurement theory, the proposed method substantially enhances the usability, credibility, and decision-support capability of metrics in real-world engineering contexts. This study thus opens a new research direction for software measurement, aligning it more closely with the epistemological standards of empirical science.

decision-makingmeasurementmetrology

Branching Out: Broadening AI Measurement and Evaluation with Measurement Trees

Sep 30, 2025
CG
Craig Greenberg
🏛️ National Institute of Standards and Technology

Current AI system evaluations suffer from fragmented assessment dimensions, heterogeneous evidence sources, and insufficient transparency. To address these challenges, this paper proposes the “Measurement Tree”—a novel multi-source fusion evaluation framework based on a hierarchical directed graph. Structured as a tree-like data model, it supports user-defined aggregation functions to unify heterogeneous metrics—including agency, business value, energy efficiency, socio-technical impact, and safety—into interpretable, multi-level representations. This work introduces, for the first time, a hierarchical graph structure as the formal output format for AI evaluation, substantially enhancing traceability and interpretability. An accompanying open-source Python library and extensive empirical validation demonstrate that the Measurement Tree improves comprehensiveness, operationality, and reproducibility in evaluating complex AI systems. It thus provides foundational infrastructure for building an open and transparent AI evaluation ecosystem.

Creating hierarchical metrics for multi-level AI system representationEnhancing transparency in AI evaluation through interpretable measurement structuresIntegrating diverse evidence types for comprehensive AI assessment

Metrics, KPIs, and Taxonomy for Data Valuation and Monetisation - Internal Processes Perspective

Dec 11, 2025
EV
Eduardo Vyhmeister
🏛️ University College Cork | Centro Tecnológico de Investigación, Desarrollo e Innovación en tecnologías de la Información y las Comunicaciones - ITI | EGI Foundation | Big Data Value Association

In data-driven economies, organizations lack systematic frameworks for evaluating and managing data value within internal business processes. To address this gap, this study develops a comprehensive data value assessment framework grounded in the Balanced Scorecard’s internal process perspective, integrating three interrelated dimensions: data quality, governance compliance, and operational efficiency. It introduces a novel, multi-layered taxonomy of data value—spanning technological, organizational, and regulatory dependencies—that resolves metric redundancy and establishes cross-dimensional conceptual linkages. Through systematic literature review, theoretical modeling, indicator clustering, and taxonomy design, the research produces a scalable, reusable data value metrics system. This system underpins standardized data valuation models and decision-support systems, offering both a methodological foundation and actionable implementation pathways for cross-sectoral data assetization. (149 words)

Develops taxonomy linking technical, organizational and regulatory indicatorsIdentifies metrics for data valuation from internal processes perspectiveLacks unified framework for measuring data value across organizations

A Task Taxonomy for Conformance Checking

Jul 16, 2025
JR
Jana-Rebecca Rehse

Existing visualization tools for compliance checking lack systematic characterization of analytical tasks, hindering rigorous effectiveness evaluation. This paper introduces the first multidimensional task taxonomy specifically designed for compliance checking, modeling core trace-to-model alignment tasks in process mining along six dimensions: objective, method, constraint type, data characteristics, data target, and cardinality. Crucially, this taxonomy explicitly links the semantic requirements of compliance checking with established visual analytics design principles—thereby bridging the semantic gap between process mining and visual analytics. It provides a reusable theoretical framework to rigorously define visualization purposes, evaluate tool effectiveness, and support co-design of analysis systems. As a result, the interpretability and practical utility of complex compliance analysis outcomes are significantly enhanced.

Clarify purposes of diverse conformance checking visualizations.Classify tasks in conformance checking analyses.Enable systematic evaluation of visualization usefulness.

Latest Papers

What's happening recently
View more

This work addresses the problem of implementation drift in evolving distributed systems, where runtime behavior gradually deviates from the original design. To tackle this issue, the paper proposes a design conformance assessment method based on distributed tracing data. It introduces, for the first time in the domain of distributed systems, conformance checking techniques from process mining, leveraging runtime traces collected via the OpenTelemetry standard and automatically comparing them against behavioral models defined at design time to quantify their alignment. The key contribution lies in establishing persistent, monitorable conformance metrics that enable continuous, automated evaluation of deviations between system implementation and design. This approach is readily applicable to modern distributed systems widely adopting OpenTelemetry for observability.

design conformancedistributed systemsimplementation drift

This study addresses the lack of systematic evaluation of data quality tools with respect to their measurement capabilities and integration with large language models (LLMs). It presents the first multidimensional assessment framework grounded in real-world enterprise use cases, systematically evaluating six prominent tools—including open-source solutions such as Great Expectations and Deequ, as well as commercial platforms like Informatica and Experian—across dimensions including rule definition, duplicate detection, metric aggregation, and uncertainty handling, along with their LLM integration mechanisms. The findings reveal that commercial tools offer more comprehensive functionality and初步 support for LLM-assisted rule generation, whereas open-source tools provide greater flexibility at the cost of higher implementation effort. Notably, none of the evaluated tools currently enable direct LLM-based data validation. This work provides empirical guidance for selecting data quality tools and advancing their integration with LLMs.

data qualitydata validationLLM integration

This study addresses the growing challenge posed by the widespread involvement of AI agents in software development, which undermines the long-standing assumption that development artifacts are exclusively produced by human professionals—an assumption underpinning traditional software metrics. The work systematically exposes how AI-generated traces compromise the foundational premises of established software measurement practices, thereby threatening the validity of prior empirical conclusions. To confront this issue, the authors propose an AI-augmented, systematic replication methodology that integrates modern data analytics with empirical software engineering techniques to rigorously re-evaluate key findings. The project advances a dynamic, reproducible, and sustainable measurement paradigm capable of adapting to evolving data ecosystems, offering a robust and timely framework for software metrics in the AI era.

AI agentsfoundational assumptionsreplication