measurement design

Designing reproducible, automated measurement and experiment pipelines that reliably quantify system behaviour (latency, energy, network effects, etc.), including instrumentation, validation on real hardware, and procedures to ensure fair comparability across services or conditions.

measurementdesign

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Reasonable Experiments in Model-Based Systems Engineering

Sep 12, 2025
JC
Johan Cederbladh
🏛️ Mälardalen University | Eindhoven University of Technology | Stellenbosch University | IT University of Copenhagen | University of Oslo | Universidade Federal Rural de Pernambuco | University of Antwerp

In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.

Deciding if existing experiments can answer new engineering questionsIntelligently reusing experiment-related data to avoid redundant experimentsManaging experimental configuration metadata and results efficiently

Technique to Baseline QE Artefact Generation Aligned to Quality Metrics

Nov 18, 2025
EF
Eitan Farchi
🏛️ IBM Research | IBM Consulting

This study addresses the uncontrolled quality of quality engineering (QE) artifacts—such as requirements specifications, test cases, and Behavior-Driven Development (BDD) scenarios—automatically generated by large language models (LLMs). We propose an iterative optimization framework integrating forward generation, backward generation, and rubric-guided scoring to enhance artifact quality along four dimensions: clarity, completeness, consistency, and testability. Our approach enables automated, quantitative, and reproducible quality assessment and improvement. Evaluated across 12 real-world projects, the method significantly improves output stability: it preserves high quality under high-quality inputs and substantially outperforms baselines under low-quality inputs. The core contribution is the first integration of backward generation with structured rubric-based guidance, establishing a closed-loop, artifact-centric quality enhancement paradigm for QE.

Ensuring generated requirements and test cases meet quality metricsEstablishing baselines for automated QE artefact quality evaluationValidating LLM outputs through reverse generation and iterative refinement

PlantD: Performance, Latency ANalysis, and Testing for Data Pipelines -- An Open Source Measurement, Testing, and Simulation Framework

Apr 14, 2025
CB
Christopher Bogart
🏛️ Carnegie Mellon University | Honda Research Institute USA, Inc.

The proliferation of sensor devices (e.g., vehicle telematics) has led to rapidly growing data pipelines, making it difficult for enterprises to quantitatively predict infrastructure costs and performance for business teams—resulting in widespread over-provisioning. Method: We propose the “Data Pipeline Wind Tunnel” paradigm, integrating synthetic workload generation, multi-dimensional metric collection (latency, throughput, resource consumption), interactive visualization, and business-hypothesis-driven “what-if” modeling for annualized cost and SLA compliance. A reusable, open-source measurement harness is implemented to support systematic pipeline benchmarking. Contribution: This work establishes, for the first time, an interpretable mapping from engineering performance metrics to business decision parameters—including annualized infrastructure cost and SLA attainment rate. Evaluated across three real-world automotive data pipelines, the framework enables cross-functional collaboration and optimization, reducing infrastructure over-provisioning by up to 42% while maintaining SLA targets.

Closing gap between technical metrics and business cost implicationsForecasting cost and performance of data pipelines for device deploymentsSimulating pipeline performance under projected real-world loads

This work addresses the problem of implementation drift in evolving distributed systems, where runtime behavior gradually deviates from the original design. To tackle this issue, the paper proposes a design conformance assessment method based on distributed tracing data. It introduces, for the first time in the domain of distributed systems, conformance checking techniques from process mining, leveraging runtime traces collected via the OpenTelemetry standard and automatically comparing them against behavioral models defined at design time to quantify their alignment. The key contribution lies in establishing persistent, monitorable conformance metrics that enable continuous, automated evaluation of deviations between system implementation and design. This approach is readily applicable to modern distributed systems widely adopting OpenTelemetry for observability.

design conformancedistributed systemsimplementation drift

Automatic Generation of Digital Twins for Network Testing

Oct 03, 2025
SD
Shenjia Ding
🏛️ University of Glasgow

Manual pre-deployment testing and validation of communication software in autonomous network evolution is time-consuming and labor-intensive. Method: This paper proposes a digital twin (DT) automated generation method aligned with the ITU-T Autonomous Networks architecture, integrating network modeling, automated orchestration, and parameter-driven simulation to generate executable, high-fidelity DT instances directly from real-world network configurations. Contribution/Results: The approach significantly reduces manual configuration overhead and enables seamless integration of the DT environment into existing verification workflows, supporting efficient execution of experimental subsystems. Experimental evaluation demonstrates that the generated DTs meet practical testing requirements in both accuracy and runtime efficiency. To the best of our knowledge, this work achieves the first end-to-end automated construction and closed-loop validation of digital twins compliant with the ITU-T G.1000 series standards.

Automating digital twin creation for network testingEnabling efficient autonomous network experimentation subsystemsReducing manual configuration effort in validation tools

Latest Papers

What's happening recently
View more

This study addresses the critical issue of declining reproducibility in quantum software defect datasets—such as Bugs4Q—due to dependency evolution, which undermines research reliability. The authors present the first systematic evaluation of this reproducibility degradation by reproducing 37 bugs across 21 Qiskit versions through 77,700 executions. Combining root cause analysis, dependency management, and API migration insights, they demonstrate that 93.6% of reproduction failures stem from environmental dependency issues rather than actual bug disappearance. Based on these findings, they propose a novel maintenance paradigm requiring source-level fixes and introduce an enhanced dataset, Bugs4Q-Robust, which boosts the reproduction rate from 16.2% to 78.4% on Qiskit v2.3.1—substantially outperforming conventional version-locking approaches.

Bugs4Qdefect datasetsdependency evolution

This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.

behavioral propertiesinterface propertiesmodel verification

This work addresses the behavioral gap between formal verification and actual execution in traditional engineering approaches, which often neglect execution semantics. To bridge this semantic divide, the paper proposes a Modeling and Simulation-Based Engineering (MSBE) methodology that explicitly treats execution semantics as a first-class engineering entity. It defines executability as the admissible model space induced by the stabilization of execution conditions and unifies model behavior with physical execution through an iterative cycle of formal execution, experimental execution, verification, and activity-mediated validation. Integrating formal methods, simulation-based verification, activity theory, and constraint modeling, MSBE establishes a general-purpose engineering framework applicable to diverse cyber-physical systems (CPS). The approach demonstrates its generality and effectiveness across four CPS categories: human-centric, biophysical, technological, and digital twin systems.

Cyber-Physical Systemsexecution semanticsformal verification

This study addresses the persistent gap between theoretical control performance and its practical realization in real-world robotic systems, often caused by inadequate discretization, insufficient real-time guarantees, and weak error handling in control software. For the first time from a software engineering perspective, the authors systematically analyze 184 open-source robotic controllers through code review, empirical analysis, and test evaluation, uncovering common deficiencies in application scenarios, implementation details, and verification practices. The findings reveal that most implementations fail to properly account for critical system constraints, and their testing strategies inadequately validate the theoretical assurances they claim. This work highlights a significant disconnect between implementation quality and theoretical promises, offering concrete directions and practical guidelines for developing reliable, verifiable robotic control software.

discretizationimplementation qualityreal-time reliability

Existing software energy measurement tools struggle to balance accuracy and overhead while often being constrained to specific hardware or programming languages, limiting their cross-platform portability. This work proposes CodeGreen, a modular energy measurement platform that innovatively integrates Tree-sitter–based AST queries to enable automatic, multi-language instrumentation. By decoupling instrumentation from measurement through an asynchronous producer-consumer architecture, CodeGreen supports fine-grained energy analysis for languages including Python, C/C++, and Java. Its Native Energy Measurement Backend (NEMB) unifies polling of hardware sensors such as Intel RAPL, NVIDIA NVML, and AMD ROCm. Evaluated on the Computer Language Benchmarks Game, CodeGreen achieves an energy estimation accuracy with a coefficient of determination of R² = 0.9934 and demonstrates near-perfect workload linearity (R² = 0.9997), offering both high precision and low overhead.

hardware couplingmeasurement accuracyportability

Hot Scholars

MH

Marco Hutter

Professor of Robotics, ETH Zurich
Legged RoboticsRoboticsControl
HB

Hyun-Bin Kim

KAIST
force torque sensorquadruped robotssensorcontrol
KS

Kyung-Soo Kim

Professor of Mechanical Engineering, KAIST
controlrobotmechatronicsmanufacturing
AE

Ahmet Emir Dirik

Bursa Uludağ University
media forensicsmultimedia forensicssignal processingpattern recognition