conduct resilience testing

Design and implement stress‑testing frameworks and synthetic‑event generators that produce scenario‑based and Monte Carlo ensembles (including spatiotemporal and Hawkes‑process simulations) to expose model and system failure modes. Construct standards‑aligned, multi‑level threat and failure taxonomies, measure and analyze distributions of adverse outcomes and robustness metrics across simulations, and translate those results into targeted mitigation and recovery actions.

conductresiliencetesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.42
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$199K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Beyond the Norm: A Survey of Synthetic Data Generation for Rare Events

Jun 04, 2025
JG
Jingyi Gu
🏛️ New Jersey Institute of Technology

Extreme events—such as stock market crashes, earthquakes, and pandemics—are rare, catastrophic, and exhibit system-wide propagation, leading to severe data scarcity that undermines data-driven modeling. To address this, we present the first systematic survey of synthetic data generation methods tailored to extreme events and propose the first dedicated generative framework for extremely rare events. We design a customized evaluation suite encompassing statistical fidelity, dependency preservation, visual plausibility, and task-oriented utility, rigorously analyzing metric validity under heavy-tailed distributions. Our framework unifies generative models (GANs, diffusion models, VAEs), large language models, statistical modeling, and targeted resampling strategies. We curate benchmark datasets across finance, meteorology, geoscience, and epidemiology, identifying underexplored domains—including behavioral finance, wildfire dynamics, and windstorm modeling—and distill key open challenges to advance the reliability and practicality of extreme-event modeling.

Addressing scarcity of training data for extreme eventsEvaluating models for heavy-tailed distributions in extreme eventsSynthetic data generation for rare extreme events

This study addresses a critical limitation of traditional financial stress testing—its reliance on subjectively designed scenarios that often overlook plausible high-risk events or introduce implausible shocks. For the first time, large deviation theory is integrated into stress testing to systematically generate extreme yet realistic stress scenarios by identifying the most likely configurations of exogenous risk factors that lead to severe losses. This approach overcomes the scarcity of historical extreme observations by combining conditional concentration analysis, probabilistic modeling of risk factor distributions, and extrapolation techniques. The method robustly reproduces stress loss distributions and key diagnostic metrics across two distinct financial network models, demonstrating consistent effectiveness even in regimes where conventional approaches completely fail.

extreme eventsfinancial risklarge deviations

Querying Labeled Time Series Data with Scenario Programs

Nov 13, 2025
EK
Edward Kim
🏛️ University of California, Berkeley | Korea University | Chalmers University of Technology | University of Gothenburg | University of California, Santa Cruz

To address the “simulation-to-reality gap”—the difficulty of reproducing simulation-identified failure scenarios in real-world autonomous driving—this paper proposes a verification method based on formal scenario modeling and time-series matching. The method formally translates abstract scenario programs written in the Scenic probabilistic programming language into computable temporal matching rules, enabling precise retrieval of failure-relevant patterns from large-scale real-world sensor data. A key contribution is the design of an efficient, linearly scalable query algorithm that supports real-time pattern matching over long temporal sequences. Experimental evaluation demonstrates that the approach achieves higher recall accuracy for critical failure scenarios than state-of-the-art commercial vision-language models, while accelerating query throughput by several orders of magnitude. This significantly improves both the efficiency and trustworthiness of transferring simulation-discovered failures to real-vehicle validation.

Bridging the sim-to-real gap in autonomous vehicle failure scenario validationDeveloping efficient algorithms to query labeled time series data matching abstract scenariosIdentifying real-world occurrences of simulated failure scenarios in sensor data

Detecting and Preventing Latent Risk Accumulation in High-Performance Software Systems

Oct 04, 2025
JA
Jahidul Arafat
🏛️ Auburn University | Oracle | Orange Business Development Limited | Bangladesh University of Professionals | Bangladesh Army International University of Science and Technology | Green University of Bangladesh

High-performance software systems accumulate latent reliability risks through aggressive optimizations; superficial performance metrics (e.g., high cache hit rates) mask underlying bottlenecks, leading to load amplification and cascading failures upon degradation. Current reliability engineering emphasizes reactive mitigation, lacking proactive identification and prevention of optimization-induced fragility. Method: We propose the first systematic framework for optimization-risk management, introducing a novel quantitative model and the Latent Risk Index (LRI). Our tripartite defense architecture—HYDRA (risk detection), RAVEN (perturbation-based validation), and APEX (risk-aware optimization)—integrates mathematical modeling, six categories of optimization-sensitive perturbation testing, and high-precision online monitoring. Contribution/Results: Experiments demonstrate 89.7% risk detection rate, >92.9% monitoring accuracy, 69.1% reduction in MTTR, annual cost savings of $1.44M, and a payback period of just 3.2 months.

Detecting latent risks in high-performance software systems with hidden vulnerabilitiesPreventing catastrophic fragility masked by exceptional performance optimizationsTransforming reliability engineering from reactive to proactive risk management

Existing power grid resilience research remains largely conceptual or focuses on isolated components, lacking a system-level, quantifiable definition and assessment framework. Method: Leveraging 15-minute-resolution customer outage time-series data and high-resolution meteorological records, we develop a spatiotemporal statistical model incorporating resilience sensitivity simulation and outage propagation dynamics inference. Contribution/Results: We propose the first system-level, empirically measurable definition of grid resilience. The model uncovers cumulative outage effects under extreme weather, inter-regional outage propagation mechanisms, and systemic response patterns. It identifies critical reinforcement nodes that reduce customer outage magnitude by nearly 50%. Validated across three major U.S. East Coast utility service territories, the model achieves high accuracy in forecasting outage progression—enabling actionable support for real-time dispatch decisions and emergency response.

Analyze large-scale outage data to understand system-level resilienceDevelop predictive model for outage progress during disastersQuantify power grid resilience against extreme weather events

Latest Papers

What's happening recently
View more

This study addresses the challenge of generating extreme joint loss scenarios and characterizing conditional distributions in multivariate heavy-tailed risk factor systems. It proposes a Self-Similar Generative Estimation (SSGEN) framework that models extremal dependence via Pareto radial components, learning from intermediate exceedances to extrapolate reliably to rarer events. The key insight is that both the conditional stress distribution and the most likely stress configuration are governed by a common limiting tail law, enabling a generative approach that ensures convergence even when the target event is absent from observed samples. The method accurately recovers rare-event probabilities and scaling laws for conditional stress scenarios, delivering a data-driven inverse stress solution with theoretical guarantees on convergence rates.

conditional lawextremal dependenceheavy-tailed risks

This study addresses the limitations of traditional risk matrices in supporting fine-grained, context-sensitive risk decision-making within complex dynamic systems. The authors propose a traceable, three-stage risk analysis framework: first, employing a multidimensional polar-coordinate heatmap to enable context-aware risk prioritization; second, constructing Bowtie causal barrier models for high-priority risks; and third, automatically transforming these Bowtie models into Bayesian networks to facilitate dynamic inference and “what-if” scenario analysis. A key innovation lies in explicitly modeling barriers as activated nodes, thereby establishing an integrated pathway from macro-level risk screening to micro-level intervention. Validation in a real-time payment gateway setting demonstrates that the proposed approach significantly enhances the transparency, auditability, and operational readiness of risk analysis.

context-sensitive triagecyber riskoperational resilience

This study addresses a critical gap in existing cybersecurity frameworks for critical infrastructure, which predominantly emphasize post-disruption recovery while lacking empirical theory on achieving limited improvements during disturbances. The work proposes the first Antifragility Theory (AFT) tailored to operational technology (OT) environments, integrating a five-state resilience system model with a mathematically rigorous definition of Jensen gain. Validation is conducted using a subset of the CISSM incident database and the HAI hardware-in-the-loop experimental platform. The research establishes two necessary conditions for antifragility verification: heterogeneous vulnerability burdens and process-level disturbance observability. Empirical findings reveal that OT systems are particularly susceptible to disruptive attacks, yet exhibit significantly reduced process-state deviations in subsequent observation windows following such attacks, indicating a measurable adaptive response inherent to the system.

antifragilitycritical infrastructurecybersecurity

This study addresses the bias in failure probability estimation caused by Markov chain samples becoming trapped in local regions when applying conventional subset simulation to multimodal failure domains. To overcome this limitation, the paper proposes a directional Subset Simulation (dSS) method that innovatively integrates directional sampling with subset simulation. By constructing nested intermediate failure domains propagated along multiple directions, dSS effectively prevents samples from being confined to localized areas at intermediate levels. The approach further combines adaptive Monte Carlo and Markov chain Monte Carlo strategies to significantly enhance the accuracy and stability of estimating probabilities of rare failure events. Numerical experiments demonstrate the superior performance of dSS in multimodal scenarios.

failure probability estimationmulti-modal failure domainsrare failure events

This study addresses the limitations of conventional subset simulation in accurately estimating failure probabilities when dealing with multiple failure regions, discontinuous, or highly nonlinear performance functions—challenges largely stemming from its reliance on the random-walk Metropolis sampler. To overcome these issues, this work proposes the first integration of the Intrepid MCMC sampler into the subset simulation framework, effectively mitigating sampling inefficiencies in multimodal, discontinuous, and high-dimensional failure domains. The proposed approach significantly enhances both the accuracy and robustness of failure probability estimation for complex reliability problems. Its superior performance is demonstrated across a range of benchmark examples spanning dimensions from 2 to 1003, confirming its effectiveness and scalability in practical applications.

discontinuous performance functionsfailure probability estimationmultimodal distributions

Hot Scholars

DW

Danny Weyns

Katholieke Universiteit Leuven
Self-adaptive systemssoftware engineeringsoftware architectureempirical methods
ZG

Zhaoya Gong

Peking University Shenzhen Graduate School
GIScienceGeoAIGeocomputationUrban and Regional Science
MH

Matthias Hollick

Professor of Computer Science, Technische Universität Darmstadt
Secure Mobile NetworkingNetwork SecurityMobile Networking
IL

Isabelle Lee

University of Southern California
MLNLPAI