perform real-world testing

Designs and executes field and real-world experiments and instrumentation to evaluate system performance, reliability, safety, and operational tradeoffs under actual deployment conditions. Builds data collection and analysis pipelines to compare methods against baselines, quantify metrics (e.g., latency, size, energy), and diagnose deployment and failure modes.

performreal-worldtesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

PlantD: Performance, Latency ANalysis, and Testing for Data Pipelines -- An Open Source Measurement, Testing, and Simulation Framework

Apr 14, 2025
CB
Christopher Bogart
🏛️ Carnegie Mellon University | Honda Research Institute USA, Inc.

The proliferation of sensor devices (e.g., vehicle telematics) has led to rapidly growing data pipelines, making it difficult for enterprises to quantitatively predict infrastructure costs and performance for business teams—resulting in widespread over-provisioning. Method: We propose the “Data Pipeline Wind Tunnel” paradigm, integrating synthetic workload generation, multi-dimensional metric collection (latency, throughput, resource consumption), interactive visualization, and business-hypothesis-driven “what-if” modeling for annualized cost and SLA compliance. A reusable, open-source measurement harness is implemented to support systematic pipeline benchmarking. Contribution: This work establishes, for the first time, an interpretable mapping from engineering performance metrics to business decision parameters—including annualized infrastructure cost and SLA attainment rate. Evaluated across three real-world automotive data pipelines, the framework enables cross-functional collaboration and optimization, reducing infrastructure over-provisioning by up to 42% while maintaining SLA targets.

Closing gap between technical metrics and business cost implicationsForecasting cost and performance of data pipelines for device deploymentsSimulating pipeline performance under projected real-world loads

This work addresses the unreliability of developer productivity dashboards, which often stems from ad hoc scripts that introduce undetected silent data gaps, eroding organizational trust. To resolve this, we propose a robust ELT pipeline grounded in DAG-based orchestration and the Medallion architecture, decoupling data extraction from transformation to preserve the immutability of raw data. Our approach introduces a state-driven dependency scheduling mechanism and, for the first time, treats metric pipelines as production-grade distributed systems. We emphasize the critical role of immutable raw history in enabling reliable metric redefinition. This methodology significantly enhances data reliability and freshness while effectively eliminating silent failures, thereby restoring organizational confidence in DevOps metrics.

Data ReliabilityDeveloper ProductivityDORA Metrics

Reasonable Experiments in Model-Based Systems Engineering

Sep 12, 2025
JC
Johan Cederbladh
🏛️ Mälardalen University | Eindhoven University of Technology | Stellenbosch University | IT University of Copenhagen | University of Oslo | Universidade Federal Rural de Pernambuco | University of Antwerp

In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.

Deciding if existing experiments can answer new engineering questionsIntelligently reusing experiment-related data to avoid redundant experimentsManaging experimental configuration metadata and results efficiently

To address the challenges of standardizing Site Reliability Engineering (SRE) practices in heterogeneous environments and balancing system reliability with development agility, this paper proposes a customizable SRE process framework. The framework integrates automated operations, multidimensional observability (metrics, logs, traces), error-budget-driven governance, standardized incident response, and progressive delivery (canary and blue-green deployments). It is designed for cross-technology-stack adaptability, enabling contextual implementation of core SRE principles. Evaluated in production systems, the framework reduced mean time to recovery by 42%, decreased unplanned outages by 67%, lowered operational staffing requirements by 35%, and achieved 99.99% service availability. Its primary contribution is the first methodology for customizing SRE processes specifically for heterogeneous environments, empirically demonstrating synergistic improvements in both system reliability and operational efficiency.

Analyzes SRE processes to boost efficiency, reduce downtimeExplores SRE for scalable, reliable software systemsPresents adaptable SRE techniques for diverse environments

Coupled Requirements-Driven Testing of CPS: From Simulation to Reality

Mar 24, 2024
AA
Ankit Agrawal
🏛️ St. Louis University | University of Innsbruck

Safety-critical small Unmanned Aircraft Systems (sUAS) lack systematic, standardized testing processes that are tightly integrated with safety analysis. Method: This paper proposes a requirement-driven coupled testing framework, introducing the novel triadic paradigm of “requirements–simulation testing–safety analysis.” It employs formal requirement modeling with bidirectional traceability, a simulation–hardware-in-the-loop cooperative testing architecture, scenario-driven test case generation, and deep integration of safety analysis methods (e.g., Fault Tree Analysis and System-Theoretic Process Analysis). Contribution/Results: Evaluated on an sUAS case study, the framework significantly improves simulation fidelity coverage and requirement coverage, enables end-to-end safety evidence generation, fills the gap in standardized sUAS testing procedures, and delivers reproducible, verifiable testing assets to support airworthiness certification.

Cyber-Physical Systems TestingSafety Analysis IntegrationStandardization

Latest Papers

What's happening recently
View more

This work addresses the challenge that existing experimental environments for distributed Cyber-Physical Systems (CPS) struggle to support reproducible, observable, and controllable integration of heterogeneous edge, fog, and cloud resources. To bridge this gap, the paper proposes a generic cloud continuum experimentation architecture grounded in the SLICES blueprint, featuring a two-layer reference model that decouples infrastructure from application logic. CPS workflows are structured along an edge–fog–cloud continuum, with deployment location, timing, and data provenance treated as core experimental dimensions. The architecture integrates virtualized and physical edge nodes, digital twin coordination, time-windowed control, and combined stream processing with cloud-side aggregation analytics, enabling multi-domain CPS applications to share programmable infrastructure and flexibly deploy and compare control and monitoring strategies. Validation through 40 systematic experiments across geographically distributed deployments—spanning renewable energy community management and AirWatch monitoring use cases—demonstrates the framework’s effectiveness and generality in hybrid physical-virtual settings.

Cloud ContinuumCyber-Physical SystemsDistributed Experimentation

This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).

Agentic SystemsMonitoringStructural Defects

This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.

Autonomous AgentsHigh Performance ComputingJob Specification Translation

Hot Scholars

DS

Davide Scaramuzza

Professor of Robotics and Perception, University of Zurich
RoboticsRobot VisionMicro Air VehiclesSLAM
JX

Jiaxu Xing

PhD Student, Robotics and Perception Group, University of Zurich
RoboticsComputer VisionMachine Learning
BZ

Boyu Zhou

Assistant Professor, SUSTech
Roboticsaerial robotsactive perceptionmobile manipulation
IK

Itzik Klein

University of Haifa
RoboticsInertial SensingData-Driven NavigationAUV
DZ

Danping Zou

Professor, Shanghai Jiao Tong University
Visual SLAMRobotic VisionVision-based navigation