Score
Designs and executes field and real-world experiments and instrumentation to evaluate system performance, reliability, safety, and operational tradeoffs under actual deployment conditions. Builds data collection and analysis pipelines to compare methods against baselines, quantify metrics (e.g., latency, size, energy), and diagnose deployment and failure modes.
The proliferation of sensor devices (e.g., vehicle telematics) has led to rapidly growing data pipelines, making it difficult for enterprises to quantitatively predict infrastructure costs and performance for business teams—resulting in widespread over-provisioning. Method: We propose the “Data Pipeline Wind Tunnel” paradigm, integrating synthetic workload generation, multi-dimensional metric collection (latency, throughput, resource consumption), interactive visualization, and business-hypothesis-driven “what-if” modeling for annualized cost and SLA compliance. A reusable, open-source measurement harness is implemented to support systematic pipeline benchmarking. Contribution: This work establishes, for the first time, an interpretable mapping from engineering performance metrics to business decision parameters—including annualized infrastructure cost and SLA attainment rate. Evaluated across three real-world automotive data pipelines, the framework enables cross-functional collaboration and optimization, reducing infrastructure over-provisioning by up to 42% while maintaining SLA targets.
This work addresses the unreliability of developer productivity dashboards, which often stems from ad hoc scripts that introduce undetected silent data gaps, eroding organizational trust. To resolve this, we propose a robust ELT pipeline grounded in DAG-based orchestration and the Medallion architecture, decoupling data extraction from transformation to preserve the immutability of raw data. Our approach introduces a state-driven dependency scheduling mechanism and, for the first time, treats metric pipelines as production-grade distributed systems. We emphasize the critical role of immutable raw history in enabling reliable metric redefinition. This methodology significantly enhances data reliability and freshness while effectively eliminating silent failures, thereby restoring organizational confidence in DevOps metrics.
In model-based systems engineering, low experimental data reuse efficiency and excessive redundant experiments hinder digital engineering agility. To address this, this paper proposes a case-based reasoning (CBR)-driven experimental management framework that explicitly integrates domain knowledge. The framework features structured experimental metadata modeling, digital twin–enabled scenario semantic alignment, and an interpretable similarity assessment mechanism to intelligently determine whether historical experiments can be transferred to address new verification queries. Its key innovation lies in embedding domain knowledge explicitly into both the CBR retrieval and adaptation stages, thereby enabling trustworthy cross-operating-condition and cross-configuration experimental data reuse. Evaluated on an industrial-scale vehicle energy system design case, the framework reduces redundant experiments by 37% and shortens early verification cycles by 42% on average, significantly enhancing iterative efficiency in digital engineering and advancing intelligent experimental management.
To address the challenges of standardizing Site Reliability Engineering (SRE) practices in heterogeneous environments and balancing system reliability with development agility, this paper proposes a customizable SRE process framework. The framework integrates automated operations, multidimensional observability (metrics, logs, traces), error-budget-driven governance, standardized incident response, and progressive delivery (canary and blue-green deployments). It is designed for cross-technology-stack adaptability, enabling contextual implementation of core SRE principles. Evaluated in production systems, the framework reduced mean time to recovery by 42%, decreased unplanned outages by 67%, lowered operational staffing requirements by 35%, and achieved 99.99% service availability. Its primary contribution is the first methodology for customizing SRE processes specifically for heterogeneous environments, empirically demonstrating synergistic improvements in both system reliability and operational efficiency.
Safety-critical small Unmanned Aircraft Systems (sUAS) lack systematic, standardized testing processes that are tightly integrated with safety analysis. Method: This paper proposes a requirement-driven coupled testing framework, introducing the novel triadic paradigm of “requirements–simulation testing–safety analysis.” It employs formal requirement modeling with bidirectional traceability, a simulation–hardware-in-the-loop cooperative testing architecture, scenario-driven test case generation, and deep integration of safety analysis methods (e.g., Fault Tree Analysis and System-Theoretic Process Analysis). Contribution/Results: Evaluated on an sUAS case study, the framework significantly improves simulation fidelity coverage and requirement coverage, enables end-to-end safety evidence generation, fills the gap in standardized sUAS testing procedures, and delivers reproducible, verifiable testing assets to support airworthiness certification.
This work addresses the challenge that existing experimental environments for distributed Cyber-Physical Systems (CPS) struggle to support reproducible, observable, and controllable integration of heterogeneous edge, fog, and cloud resources. To bridge this gap, the paper proposes a generic cloud continuum experimentation architecture grounded in the SLICES blueprint, featuring a two-layer reference model that decouples infrastructure from application logic. CPS workflows are structured along an edge–fog–cloud continuum, with deployment location, timing, and data provenance treated as core experimental dimensions. The architecture integrates virtualized and physical edge nodes, digital twin coordination, time-windowed control, and combined stream processing with cloud-side aggregation analytics, enabling multi-domain CPS applications to share programmable infrastructure and flexibly deploy and compare control and monitoring strategies. Validation through 40 systematic experiments across geographically distributed deployments—spanning renewable energy community management and AirWatch monitoring use cases—demonstrates the framework’s effectiveness and generality in hybrid physical-virtual settings.
This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.