Score
Specifying and running mechanical and user-centered tests and metrics (e.g., repeatability, force output, wear, comfort) to evaluate long-term reliability, robustness, and integration of devices or sensors.
The PHM (Prognostics and Health Management) community has long suffered from a lack of systematically evaluated, freely accessible degradation datasets. Method: This work establishes the first multi-dimensional unified evaluation framework for PHM datasets, incorporating critical dimensions—data provenance, equipment types, sensor configurations, failure modes, and annotation completeness—integrated with structured metadata analysis, cross-dataset comparative assessment, task-specific PHM mapping, and physics-of-failure-informed semantic annotation. Contribution/Results: We systematically curate and analyze 32 high-quality public datasets, identifying 11 recurrent deficiencies. Based on this analysis, we provide task-oriented data selection guidelines and benchmarking recommendations. This study fills a critical gap in systematic surveys of PHM public data resources, explicitly delineates the applicability boundaries and modeling limitations of existing datasets, and has been widely cited and adopted within the PHM research community.
Statistical Query (SQ) algorithms—such as Monte Carlo and importance sampling—exhibit poor reproducibility in standardized robotic performance testing. Method: This paper proposes the first lightweight, parameterized, and adaptive general correction framework that provides provable reproducibility guarantees for arbitrary SQ algorithms, without reliance on specific hardware or operational procedures. The framework integrates statistical query theory, adaptive sampling, and rigorous error-bound analysis to jointly optimize accuracy and efficiency. Contribution/Results: Evaluated across three canonical robotic domains—industrial manipulators, autonomous driving risk assessment, and humanoid motion control—the framework improves test reproducibility by 92%, reduces cross-platform variance by 76%, and achieves 99.3% reproducibility in instruction-tracking evaluation—significantly surpassing conventional reproducibility enhancement paradigms.
Existing mission profile modeling approaches for electric and autonomous vehicles provide only aggregated histograms of single stress parameters (e.g., temperature or voltage), lacking multidimensional coupling, temporal evolution, and full-lifecycle characterization. To address this, we propose a novel mission profile modeling framework grounded in functional temporal modeling, which jointly captures the dynamic evolution of multiple stress parameters—including temperature, humidity, and voltage. The framework supports configurable time-granularity sampling and user-defined quantile analysis, and embeds anomaly detection and data integrity protection mechanisms to ensure compliant, trustworthy data sharing across the supply chain (suppliers–OEMs–end users). For the first time, our framework enables high-fidelity, quantifiable, multi-stress-coupled mission profile modeling with built-in security and collaborative capabilities.
This work addresses the limitations of traditional Failure Modes, Effects, and Diagnostic Analysis (FMEDA) in automotive ASIC functional safety verification, where expert judgment is used to estimate failure mode distributions and diagnostic coverage without quantifying associated uncertainties, thereby compromising reliability. For the first time, error propagation theory is systematically integrated into FMEDA to construct uncertainty models for both failure mode distributions and diagnostic coverage. This enables quantitative computation of the maximum deviations and confidence intervals for the Single-Point Fault Metric (SPFM) and Latent Fault Metric (LFM). Furthermore, an Error Importance Indicator (EII) is introduced to trace the key contributors driving overall uncertainty. The proposed approach significantly enhances the transparency and credibility of FMEDA, offering a scientifically rigorous and quantifiable foundation for compliance with ISO 26262.
This work proposes a novel approach to black-box testing of Functional Mock-up Units (FMUs) by integrating large language models (LLMs) with a human-in-the-loop mechanism. Addressing the inefficiency and poor interpretability of traditional FMU-based dynamic simulation testing—which relies on manually crafted scenarios—the method automatically generates structured Given-When-Then test objectives from FMU interface and functional specifications, and constructs complete test plans comprising input sequences and assertion oracles. Upon simulation execution, the framework produces visualizable logs and statistical evaluation metrics. The approach significantly enhances test design efficiency and result interpretability, facilitates test asset reuse, and demonstrates effectiveness on a lubricating oil cooling system by autonomously generating executable test scenarios and delivering objective-level pass-rate analysis.
Safety-critical small Unmanned Aircraft Systems (sUAS) lack systematic, standardized testing processes that are tightly integrated with safety analysis. Method: This paper proposes a requirement-driven coupled testing framework, introducing the novel triadic paradigm of “requirements–simulation testing–safety analysis.” It employs formal requirement modeling with bidirectional traceability, a simulation–hardware-in-the-loop cooperative testing architecture, scenario-driven test case generation, and deep integration of safety analysis methods (e.g., Fault Tree Analysis and System-Theoretic Process Analysis). Contribution/Results: Evaluated on an sUAS case study, the framework significantly improves simulation fidelity coverage and requirement coverage, enables end-to-end safety evidence generation, fills the gap in standardized sUAS testing procedures, and delivers reproducible, verifiable testing assets to support airworthiness certification.
This work addresses the challenge that human-centered AI systems, when deployed over extended periods, often violate safety and sustainability requirements when encountering unforeseen edge cases. To overcome the limitations of conventional testing approaches—particularly in terms of resource constraints, computational demands, and complex human–machine interactions—the study proposes a novel verification and design framework that integrates model-driven engineering, personalized modeling, and AI safety analysis. By establishing a personalized, model-driven analytical mechanism, the approach effectively identifies and mitigates latent safety and sustainability risks emerging during long-term deployment. This significantly enhances the robustness, reliability, and trustworthiness of human-centered cyber-physical systems operating in real-world environments.
This study addresses the challenge of accurately assessing the reuse feasibility of returned products in circular manufacturing, where heterogeneous conditions impede reliable evaluation. The authors propose an integrated reliability workflow that synergistically combines functional behavior prediction with material fatigue analysis. Specifically, a convolutional encoder extracts force-torque loading patterns, while an LSTM network predicts Gaussian distributions for nine functional variables. Concurrently, finite element stress reconstruction is coupled with S–N/Miner’s rule (augmented by Haibach correction) and Paris’ law to evaluate output shaft fatigue. A streaming replay algorithm then fuses functional, material, and system-level reliability trajectories. This approach represents the first uncertainty-aware co-integration of functional forecasting and component-level fatigue assessment. Experimental results demonstrate an average accuracy of 0.9652 for the nine output variables within a 2% tolerance, with R² values of 0.9750 and 0.9924 for drive current and rotational speed, respectively, confirming its efficacy and reliability calibration capability for precision reuse decisions.
This study addresses safety risks arising from software faults in cyber-physical systems (CPS) for electric bicycles. We propose a simulation-driven functional Failure Mode and Effects Analysis (FMEA) method, leveraging Simulink Fault Analyzer to construct fault models, integrated with expert review and a systematic FMEA process to close the loop among fault modeling, simulation-based analysis, and impact assessment. Experimental evaluation identified 13 real-world faults with 100% model accuracy; among them, five revealed previously unrecognized safety implications, and 38.4% induced anomalous system behavior. The study distills ten reusable engineering practice guidelines, significantly enhancing the effectiveness and practicality of FMEA in industrial-scale CPS. It provides empirical validation and methodological contributions toward the operational deployment of simulation-driven safety analysis.
This study addresses the challenge of simultaneously achieving accurate point estimates and reliable prediction intervals for bearing remaining useful life (RUL) under varying operating conditions, where conventional evaluation metrics often obscure conditional failure modes. The authors propose a predictive representation model grounded in empirical residual calibration, integrating vibration signals, engineered features, and operational conditions. A bounded reliability assessment protocol is introduced, enabling condition-stratified diagnostics on rigorously partitioned datasets. The analysis reveals substantially degraded coverage—dropping as low as 0.666—under specific conditions such as low load or high rotational speed, and identifies missing raw sensor channels as a critical failure mode. Following residual calibration, the method achieves a 90% prediction interval coverage of 0.941, with a normalized mean absolute error (NMAE) of 0.1477 for the primary model.
This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).