performance metrics

Designs, implements, and analyzes quantitative measures used to evaluate how well systems, models, algorithms, or processes meet specified objectives; this includes defining metric formulas, computing and visualizing results, assessing statistical validity and sensitivity, and selecting metrics that capture relevant trade-offs and align with evaluation goals.

performancemetrics

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.61
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$194K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing software engineering metrics often fail to effectively support critical development decisions—such as whether refactoring is necessary or whether testing is sufficient—thereby limiting their practical utility. This work addresses this gap by systematically introducing metrological principles (the science of measurement) into the domain of software measurement for the first time. It proposes a metrology-informed approach to metric modeling and evaluation, establishing a rigorous scientific foundation for the design of software metrics. By grounding metric development in established measurement theory, the proposed method substantially enhances the usability, credibility, and decision-support capability of metrics in real-world engineering contexts. This study thus opens a new research direction for software measurement, aligning it more closely with the epistemological standards of empirical science.

decision-makingmeasurementmetrology

Root/Additional Metric (RoAM) framework: a guide for goal-centred metric construction

Jul 02, 2025
LE
Luke E. B. Goodyear
🏛️ Queen’s University Belfast

Existing performance measurement frameworks struggle to simultaneously satisfy customizability, interpretability, and mathematical tractability in interdisciplinary contexts. Method: This paper proposes a goal-oriented, customizable metric construction framework featuring a novel “base metric–auxiliary metric” dichotomy. Integrating utility theory and multi-criteria decision analysis, it introduces an uncertainty-aware utility function and establishes a systematic metric decomposition–synthesis workflow. Contributions: (1) It reduces reliance on complex mathematical formalisms, enhancing applicability under resource constraints or high uncertainty; (2) it ensures metric transparency, traceability, and domain adaptability; and (3) it enables quantitative assessment of goal attainment, real-time progress monitoring, and downstream statistical modeling and decision optimization. The framework has been empirically validated across diverse disciplines, demonstrating generality and extensibility.

Combines decision analysis and utility theory to quantify goal achievementDevelops a framework for constructing customizable performance metrics across disciplinesDivides criteria into root and additional groups for flexible metric design

This work addresses the problem of implementation drift in evolving distributed systems, where runtime behavior gradually deviates from the original design. To tackle this issue, the paper proposes a design conformance assessment method based on distributed tracing data. It introduces, for the first time in the domain of distributed systems, conformance checking techniques from process mining, leveraging runtime traces collected via the OpenTelemetry standard and automatically comparing them against behavioral models defined at design time to quantify their alignment. The key contribution lies in establishing persistent, monitorable conformance metrics that enable continuous, automated evaluation of deviations between system implementation and design. This approach is readily applicable to modern distributed systems widely adopting OpenTelemetry for observability.

design conformancedistributed systemsimplementation drift

Branching Out: Broadening AI Measurement and Evaluation with Measurement Trees

Sep 30, 2025
CG
Craig Greenberg
🏛️ National Institute of Standards and Technology

Current AI system evaluations suffer from fragmented assessment dimensions, heterogeneous evidence sources, and insufficient transparency. To address these challenges, this paper proposes the “Measurement Tree”—a novel multi-source fusion evaluation framework based on a hierarchical directed graph. Structured as a tree-like data model, it supports user-defined aggregation functions to unify heterogeneous metrics—including agency, business value, energy efficiency, socio-technical impact, and safety—into interpretable, multi-level representations. This work introduces, for the first time, a hierarchical graph structure as the formal output format for AI evaluation, substantially enhancing traceability and interpretability. An accompanying open-source Python library and extensive empirical validation demonstrate that the Measurement Tree improves comprehensiveness, operationality, and reproducibility in evaluating complex AI systems. It thus provides foundational infrastructure for building an open and transparent AI evaluation ecosystem.

Creating hierarchical metrics for multi-level AI system representationEnhancing transparency in AI evaluation through interpretable measurement structuresIntegrating diverse evidence types for comprehensive AI assessment

Towards Measurement Theory for Artificial Intelligence

Jul 07, 2025
EP
Elija Perrier
🏛️ UTS

Current AI capability evaluation lacks a formal measurement foundation, suffers from cross-system and cross-method incomparability, and remains disconnected from quantitative risk analysis in engineering safety. To address these issues, this paper proposes a hierarchical AI measurement theory framework that rigorously distinguishes between direct and indirect observables and formally characterizes how AI capability definitions depend on specific measurement operations and scales. Methodologically, the framework integrates classical measurement theory, formal modeling, and quantitative risk analysis techniques, drawing upon established paradigms from engineering and safety science. Its core contribution is the first systematic, calibratable, and traceable taxonomy of AI phenomena and capability representations. This enables standardized, reproducible AI system evaluation—significantly enhancing the reliability and interoperability of assessment outcomes across scientific validation, engineering deployment, and regulatory decision-making.

Compare AI systems and evaluation methods using standardized metricsDevelop a formal theory for measuring artificial intelligence capabilitiesIntegrate AI assessments with engineering risk analysis techniques

Latest Papers

What's happening recently
View more

Frameworks such as SPACE, DevEx, and DORA established that developer productivity is inherently multidimensional, but left practitioners with a practical question: what should we measure, and how should we use it to improve? This paper introduces Engineering Thrive (EngThrive), a measurement and improvement system developed and deployed across Microsoft's engineering organization. EngThrive organizes productivity around three dimensions - Speed, Ease, and Quality - with Thriving as a guardrail to ensure developer wellbeing improves alongside performance. Within each dimension, outcome-oriented North Star metrics are paired with diagnostic submetrics, combining system telemetry with developer surveys to provide both scale and context. We describe the design principles that guide metric selection, including an approach in which well-chosen metrics align "gaming" behavior with genuine improvement. We also outline the data platform, survey program, and dashboard ecosystem required to operationalize this approach in practice, and present case studies demonstrating how outcome-oriented measurement enables sustained, system-level improvements. Finally, we show that EngThrive functions as a general-purpose evaluation language, applicable not only to developer tools and AI, but to organizational policies, work environments, and other factors that shape how developers experience their work. We offer EngThrive as a concrete model for organizations seeking to move beyond measuring activity toward improving outcomes.

developer experiencedeveloper productivityengineering effectiveness

Technique to Baseline QE Artefact Generation Aligned to Quality Metrics

Nov 18, 2025
EF
Eitan Farchi
🏛️ IBM Research | IBM Consulting

This study addresses the uncontrolled quality of quality engineering (QE) artifacts—such as requirements specifications, test cases, and Behavior-Driven Development (BDD) scenarios—automatically generated by large language models (LLMs). We propose an iterative optimization framework integrating forward generation, backward generation, and rubric-guided scoring to enhance artifact quality along four dimensions: clarity, completeness, consistency, and testability. Our approach enables automated, quantitative, and reproducible quality assessment and improvement. Evaluated across 12 real-world projects, the method significantly improves output stability: it preserves high quality under high-quality inputs and substantially outperforms baselines under low-quality inputs. The core contribution is the first integration of backward generation with structured rubric-based guidance, establishing a closed-loop, artifact-centric quality enhancement paradigm for QE.

Ensuring generated requirements and test cases meet quality metricsEstablishing baselines for automated QE artefact quality evaluationValidating LLM outputs through reverse generation and iterative refinement

Metrics, KPIs, and Taxonomy for Data Valuation and Monetisation - Internal Processes Perspective

Dec 11, 2025
EV
Eduardo Vyhmeister
🏛️ University College Cork | Centro Tecnológico de Investigación, Desarrollo e Innovación en tecnologías de la Información y las Comunicaciones - ITI | EGI Foundation | Big Data Value Association

In data-driven economies, organizations lack systematic frameworks for evaluating and managing data value within internal business processes. To address this gap, this study develops a comprehensive data value assessment framework grounded in the Balanced Scorecard’s internal process perspective, integrating three interrelated dimensions: data quality, governance compliance, and operational efficiency. It introduces a novel, multi-layered taxonomy of data value—spanning technological, organizational, and regulatory dependencies—that resolves metric redundancy and establishes cross-dimensional conceptual linkages. Through systematic literature review, theoretical modeling, indicator clustering, and taxonomy design, the research produces a scalable, reusable data value metrics system. This system underpins standardized data valuation models and decision-support systems, offering both a methodological foundation and actionable implementation pathways for cross-sectoral data assetization. (149 words)

Develops taxonomy linking technical, organizational and regulatory indicatorsIdentifies metrics for data valuation from internal processes perspectiveLacks unified framework for measuring data value across organizations

Facing persistent declines in customer satisfaction among small- and medium-sized enterprises (SMEs) in the IT services sector—and intensifying competitive pressure from rivals delivering superior service at lower costs—this study proposes an integrated service quality assessment framework combining Lean Six Sigma and the SERVQUAL model. Methodologically, it systematically unifies the DMAIC methodology, SERVQUAL’s five-dimensional gap analysis, and Six Sigma’s data-driven statistical process control techniques to establish a quantifiable, traceable pathway for service improvement. Empirical validation demonstrates that the framework precisely identifies five root causes of customer dissatisfaction, increases customer satisfaction significantly, and reduces customer acquisition cost by 18.3%. This work bridges a critical methodological gap in service quality management: the absence of a quantitatively rigorous, end-to-end integrated approach—from diagnostic assessment to closed-loop optimization. It offers SME IT service providers a novel quality management paradigm that balances theoretical rigor with practical implementability.

Addresses the lack of robust service quality measurement systems causing customer attritionIntegrates Lean Six Sigma with SERVQUAL to identify root causes of dissatisfactionProposes a framework to measure customer satisfaction in computer service companies

This work addresses the limitations of existing formalisms for hyperproperties in capturing quantitative aspects inherent in real-world systems, such as numerical relationships in information flow control. To overcome this, the paper introduces Quantitative Hyper-Logic (QHL), a novel framework that reformulates hyperproperty specifications using measure theory, replacing classical Boolean quantifiers with measures to support nested quantitative structures. Leveraging Hoeffding’s inequality and extreme value theory, the authors develop an efficient statistical verification algorithm and provide rigorous analyses of sample complexity and statistical guarantees. Experimental evaluation on quantitative information-flow benchmarks demonstrates that QHL substantially outperforms conventional qualitative approaches, offering superior expressiveness and verification capabilities that better align with the demands of practical systems.

hyperpropertiesinformation flow controlmeasure-based quantification

Hot Scholars

SY

Samuel Yen-Chi Chen

Wells Fargo
quantum computationquantum informationmachine learningquantum machine learning
DT

Dzmitry Tsetserukou

Associate Professor, Skolkovo Institute of Science and Technology (Skoltech)
RoboticsHapticsUAV SwarmAI
IS

Ilia Sucholutsky

New York University
deep learningrepresentation learningsmall datarepresentational alignment
KM

Katherine M. Collins

Machine Learning PhD Student at the University of Cambridge
Cognitive ScienceMachine LearningBayesian StatisticsHuman-AI Interaction