robustness and failure-mode analysis

Designs and performs analyses and tests to discover how systems, models, or components behave and break under stress, perturbations, faults, or adversarial inputs, producing failure-mode catalogs, root-cause analyses, and quantitative robustness metrics. Uses those results to specify and implement mitigations, fault-tolerant designs, and verification or monitoring procedures that reduce sensitivity to identified failure modes and preserve acceptable behavior under expected and unexpected conditions.

robustnessandfailure-modeanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.47
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$207K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Failure Modes and Effects Analysis: An Experience from the E-Bike Domain

Sep 19, 2025
AB
Andrea Bombarda
🏛️ University of Bergamo | McMaster University

This study addresses safety risks arising from software faults in cyber-physical systems (CPS) for electric bicycles. We propose a simulation-driven functional Failure Mode and Effects Analysis (FMEA) method, leveraging Simulink Fault Analyzer to construct fault models, integrated with expert review and a systematic FMEA process to close the loop among fault modeling, simulation-based analysis, and impact assessment. Experimental evaluation identified 13 real-world faults with 100% model accuracy; among them, five revealed previously unrecognized safety implications, and 38.4% induced anomalous system behavior. The study distills ten reusable engineering practice guidelines, significantly enhancing the effectiveness and practicality of FMEA in industrial-scale CPS. It provides empirical validation and methodological contributions toward the operational deployment of simulation-driven safety analysis.

Evaluating simulation-driven FMEA effectiveness in e-Bike safety analysisModeling 13 realistic faults to detect CPS safety breachesValidating model accuracy and fault impact through expert feedback

Quantitative Measurement of Cyber Resilience: Modeling and Experimentation

Mar 28, 2023
MJ
Michael J. Weisman
🏛️ DEVCOM Army Research Laboratory | Pennsylvania State University | ICF International | University of California, Irvine

Current cyber-physical systems (CPS) in vehicular environments lack quantitative, experimentally grounded methods for assessing network resilience. Method: This study constructs an experimental testbed replicating real-world truck operational conditions and conducts multiple rounds of malware injection attacks, simultaneously collecting network- and physical-layer data on resistance and recovery behaviors. Contribution/Results: We introduce the novel concept of “bonware” to holistically characterize both cybersecurity defense capability and physical resilience, formalized via an analytically tractable mathematical model. We further define and extract experimentally identifiable, quantitative resilience metrics—termed elastic features—for the first time. Sensitivity analysis confirms these metrics exhibit significant discriminability with respect to attack intensity, defensive strategies, and physical redundancy. This work bridges a critical gap by advancing vehicular CPS resilience from qualitative description to quantifiable, comparable, and optimizable measurement.

Attack RecoveryCyber ResilienceMeasurement Tools

From PREVENTion to REACTion: Enhancing Failure Resolution in Naval Systems

Aug 21, 2025
MT
Maria Teresa Rossi
🏛️ University of Milano -Bicocca

Naval systems frequently exhibit anomalous behaviors due to wear, misuse, or component failures—challenges that hinder timely detection and precise remediation. To address this, we propose a predictive-diagnostic closed-loop framework that tightly integrates the existing failure prediction system PREVENT with a newly designed responsive troubleshooting module, REACT. Methodologically, the framework synergizes multi-source time-series anomaly detection with domain-knowledge-driven fault-isolation process modeling, enabling end-to-end automation—from anomaly alerting and root-cause localization to actionable remediation recommendations. Evaluated on operational shipboard systems deployed by Fincantieri, the framework reduces mean time to fault localization by 42%, significantly improves operational response efficiency, and demonstrates strong generalizability across diverse industrial domains.

Enhancing failure detection and resolution in naval systemsExtending predictive maintenance to industrial productsIntegrating anomaly detection with troubleshooting procedures

Automated Statistical Testing and Certification of a Reliable Model-Coupling Server for Scientific Computing

May 14, 2025
SW
Seth Wolfgang
🏛️ Indiana University | Ball State University

This work addresses the challenge of reliability verification for web services coupling multi-physics, multi-scale models in scientific computing. We propose a novel approach integrating serialized formal specifications with usage-driven statistical testing. Our method comprises constructing executable specifications, modeling realistic usage scenarios, performing adaptive statistical testing, and estimating reliability confidence—thereby overcoming the limited coverage of conventional unit testing. To our knowledge, this is the first framework that synergistically combines formal specification and statistical testing for certification of coupled services, introducing quantifiable reliability metrics. Empirical evaluation demonstrates that the method effectively uncovers critical failure paths missed by unit testing and achieves a certified reliability level of ≥0.999 for coupled service controllers at a 95% confidence level.

Certifying robustness with quantitative reliability statisticsDetecting code failures via statistical testing methodsEnsuring reliability of model-coupling web services

Distributed systems operating in complex environments are highly susceptible to failures and adversarial behaviors, making their performance difficult to predict directly from formal designs. This work proposes the first framework that systematically enables performance prediction from formal models by integrating a reusable fault-injection library with a modeling methodology based on the Maude language. The authors develop an automated tool, PERF, which combines model composition and statistical analysis techniques to accurately estimate key performance metrics—such as throughput and latency—across diverse failure scenarios. Experimental evaluation demonstrates that PERF’s predictions align closely with measurements from real-world deployments on representative distributed systems, significantly enhancing the practical utility of formal methods in performance assessment.

adversarial behaviorsdistributed systemsfaults

Latest Papers

What's happening recently
View more

This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).

Agentic SystemsMonitoringStructural Defects

This study addresses the challenge of efficiently localizing faulty modules in automotive system-level 0D simulations following model updates—a process that traditionally incurs high verification costs and prolonged cycles. To overcome this, the authors propose a novel diagnostic approach based on graph-structured modeling, which uniquely integrates Dynamic Mode Decomposition (DMD), linear programming, and autoencoders to embed system simulation behaviors into a graph representation. This framework enables automatic fault module identification with only a minimal number of simulation runs. The proposed method substantially reduces computational overhead, enhances fault localization efficiency, and seamlessly integrates into existing engineering validation workflows, offering both practical utility and strong scalability.

0D modelsfault detectionmodel updating

This work addresses the limitation of traditional Failure Modes and Effects Analysis (FMEA) in automotive semiconductors, which focuses solely on functional safety while neglecting the synergistic vulnerabilities and common-cause consequences arising from interactions with cybersecurity. To bridge this gap, the authors propose a unified Functional Safety and Cybersecurity Threat and Risk Analysis (FTMEA) framework that introduces, for the first time, quantifiable Cross-Domain Correlation Factors (CDCFs). These CDCFs integrate expert knowledge, static structural analysis (e.g., controllability and observability), and empirical data from fault and attack injection experiments to enable a cohesive risk modeling and prioritization mechanism. Applied to an automotive ASIC configuration register case study, the approach successfully identifies cross-domain risks overlooked by conventional FMEA and TARA, significantly enhancing the effectiveness of mitigation strategies and providing traceable, quantifiable risk assessment evidence.

automotive semiconductorscross-domain correlationcybersecurity

Querying Labeled Time Series Data with Scenario Programs

Nov 13, 2025
EK
Edward Kim
🏛️ University of California, Berkeley | Korea University | Chalmers University of Technology | University of Gothenburg | University of California, Santa Cruz

To address the “simulation-to-reality gap”—the difficulty of reproducing simulation-identified failure scenarios in real-world autonomous driving—this paper proposes a verification method based on formal scenario modeling and time-series matching. The method formally translates abstract scenario programs written in the Scenic probabilistic programming language into computable temporal matching rules, enabling precise retrieval of failure-relevant patterns from large-scale real-world sensor data. A key contribution is the design of an efficient, linearly scalable query algorithm that supports real-time pattern matching over long temporal sequences. Experimental evaluation demonstrates that the approach achieves higher recall accuracy for critical failure scenarios than state-of-the-art commercial vision-language models, while accelerating query throughput by several orders of magnitude. This significantly improves both the efficiency and trustworthiness of transferring simulation-discovered failures to real-vehicle validation.

Bridging the sim-to-real gap in autonomous vehicle failure scenario validationDeveloping efficient algorithms to query labeled time series data matching abstract scenariosIdentifying real-world occurrences of simulated failure scenarios in sensor data

Hot Scholars

YF

Yihe Fan

Unknown affiliation
AI safety
MY

Min Yang

Bytedance
Vision Language ModelComputer VisionVideo Understanding
JD

Jiarun Dai

Assistant Professor, Fudan Univerisity
Vulnerability DetectionAI System Security
YS

Yizhou Sun

Professor, Computer Science, UCLA
Information NetworksKnowledge GraphsGraph Neural NetworksData Mining