resilient systems design

Designs, builds, and analyzes system architectures, processes, and requirements to ensure they withstand, absorb, and rapidly recover from disruptions; this includes identifying minimum viable recovery capabilities, prioritizing capability-centric recovery design, restructuring architecture for rapid recovery, and trading off restoration scope against cost.

resilientsystemsdesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.64
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$210K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Modeling and Simulation of Data Protection Systems for Business Continuity and Disaster Recovery

Dec 01, 2025
SN
Sašo Nikolovski
🏛️ AUE -FON University | University "St. Kliment Ohridski"

In cloud environments, selecting optimal data protection strategies for business continuity and disaster recovery remains challenging due to the lack of quantitative foundations for evaluating reliability and aligning with organizational Recovery Time Objectives (RTOs) and operational requirements. Method: This paper proposes an integrated assessment framework that synergistically combines system dynamics modeling and simulation-based optimization. It quantitatively evaluates key performance indicators—including recovery timeliness, data integrity, and system robustness—across public and hybrid cloud scenarios by simulating mainstream recovery mechanisms. Contribution/Results: The framework innovatively applies system dynamics to model time-varying dependencies during recovery processes and establishes interpretable, traceable mappings between policy parameters, technical metrics, and business objectives. Empirical validation demonstrates its reproducibility and practical utility, providing cloud-native organizations with a quantifiable, verifiable, and actionable decision-support methodology for data protection strategy selection.

Comparative analysis of cloud-based recovery solutions for reliabilityModeling and simulation of data protection systems for business continuityProposes a framework for selecting and maintaining organizational recovery solutions

This study addresses the critical challenge that restoring IT backups alone is insufficient to resume production after ransomware attacks on manufacturing systems, due to deep interdependencies among IT, operational technology (OT), physical processes, identity management, and supply chains. The work reframes recovery as a problem of interdependent continuity in critical infrastructure and introduces, for the first time, the concept of “Minimum Viable Factory Recovery” (MVF Recovery), shifting the objective from full-system restoration to capability-oriented minimal trusted operations. Drawing on a PRISMA-guided multi-source systematic review integrating academic literature, standards, government guidelines, and real-world incidents, the study identifies nine failure modes in recovery efforts and proposes a capability-centered recovery framework. It further establishes an evidence-driven recovery lifecycle model and outlines directions for benchmarking, offering actionable recovery targets for critical manufacturing infrastructure.

critical infrastructuremanufacturing systemsMinimum Viable Factory

This work proposes a systematic approach to derive task effectiveness requirements in the absence of explicit user needs. The method deconstructs task intent into context, functionality, constraints, critical dimensions, performance attributes, and architectural solutions, and introduces a task complexity factor to quantify the impact of external challenges and technology maturity. By integrating Best-Worst Scaling, it prioritizes critical dimensions based on stakeholder judgments. Through task decomposition modeling and quantitative complexity analysis, the framework supports integration with UAF/SysML artifacts and establishes a traceable mechanism for generating Tier 1 and Tier 2 requirements. The approach is validated using a close air support mission case study, effectively addressing a critical gap in requirements engineering when clear initial inputs are unavailable.

adaptive methodmission complexitymission effectiveness

Enhancing Energy Sector Resilience: Integrating Security by Design Principles

Feb 18, 2024
DS
Dov Shirtz
🏛️ Ben-Gurion University | Shamoon College of Engineering

Energy-sector industrial control systems (ICS) exhibit insufficient security resilience and overreliance on reactive, post-incident remediation. Method: This paper proposes a layered, implementable Security-by-Design (SbD) framework and a deployable set of security requirements tailored to critical infrastructure. Integrating systems engineering, ICS-specific security architecture, organizational behavior principles, and continuous monitoring, the approach spans the entire lifecycle—design, development, deployment, and operations—while ensuring alignment with IEC 62443 and NIST SP 800-82. Contribution/Results: It represents the first systematic, end-to-end operationalization of SbD in energy ICS contexts, enabling a paradigm shift from passive incident response to inherent, “native immunity.” The resulting scalable, auditable, and standards-coordinated SbD implementation guide supports the development of high-assurance, resilient, and sustainably evolvable cybersecurity ecosystems.

Enhancing energy sector resilience through Security by Design (SbD) principlesEstablishing an SbD-driven ecosystem to combat cyber threats effectivelyIntegrating SbD in industrial control systems lifecycle for robust security

Quantitative Measurement of Cyber Resilience: Modeling and Experimentation

Mar 28, 2023
MJ
Michael J. Weisman
🏛️ DEVCOM Army Research Laboratory | Pennsylvania State University | ICF International | University of California, Irvine

Current cyber-physical systems (CPS) in vehicular environments lack quantitative, experimentally grounded methods for assessing network resilience. Method: This study constructs an experimental testbed replicating real-world truck operational conditions and conducts multiple rounds of malware injection attacks, simultaneously collecting network- and physical-layer data on resistance and recovery behaviors. Contribution/Results: We introduce the novel concept of “bonware” to holistically characterize both cybersecurity defense capability and physical resilience, formalized via an analytically tractable mathematical model. We further define and extract experimentally identifiable, quantitative resilience metrics—termed elastic features—for the first time. Sensitivity analysis confirms these metrics exhibit significant discriminability with respect to attack intensity, defensive strategies, and physical redundancy. This work bridges a critical gap by advancing vehicular CPS resilience from qualitative description to quantifiable, comparable, and optimizable measurement.

Attack RecoveryCyber ResilienceMeasurement Tools

Latest Papers

What's happening recently
View more

This study addresses the inadequacy of current IT compliance–oriented cybersecurity policies in safeguarding the physical safety of cyber-physical systems, as digital failures often precipitate real-world harm. By coding 292 critical infrastructure policies (2000–2025) and aligning them with the NIST SP 800-160 Vol. 2 resilience lifecycle, the research reveals a significant misalignment between prevailing policy approaches—overreliant on IT control catalogs during resistance and recovery phases—and actual physical risks. The work proposes a modernized “duty of reasonable care” standard centered on hazard-specific traceability, structured assurance cases, and cyber resilience engineering. It identifies three critical disconnects: misaligned delegation of standards, reduction of recovery mechanisms to mere incident reporting, and uneven sectoral adaptability. The study further outlines a viable pathway for federal policy that integrates engineering implementation with targeted incentives.

critical infrastructurecyber safetycyber-physical systems

This study addresses the vulnerability of complex systems to self-induced collapse under shocks, a challenge inadequately met by conventional approaches that struggle to balance robustness and adaptability. The work distinguishes two dynamic regimes—phase-separated systems and highly volatile, interwoven systems—and conceptualizes resilience as an emergent property arising from multi-agent interactions. Rather than advocating mere restoration to prior states, it proposes systemic transformation to enhance recovery capacity. Methodologically, the research integrates data-driven multi-agent modeling, knowledge graphs, and artificial intelligence tools. Large-scale simulations reveal that optimizing for peak performance often undermines resilience, whereas second-order interventions leveraging positive feedback mechanisms can effectively reconfigure system architecture, thereby substantially strengthening overall resilience.

agent-based modelingbreakdownrecovery

This study addresses the challenge faced by production system engineers in automatically verifying production line layouts due to limited knowledge of PDDL and planning theory. To bridge this gap, the authors propose a novel approach based on an Asset Administration Shell (AAS) capability model that natively generates complete PDDL planning problems directly from domain-level descriptions, eliminating the need for PDDL-specific submodels. The method integrates four Industry 4.0 standards—VDI 3682, IEC 61360-1, IDTA 02011, and IDTA 02016—to construct the AAS and employs an extraction algorithm to automatically translate multi-AAS architectures into PDDL domains. In a laboratory case study, the approach enabled engineers to systematically compare four layout variants by modifying only the AAS model, significantly lowering the barrier to adopting automated planning in industrial settings.

Asset Administration ShellAutomated PlanningCapability Modeling

This work addresses the inefficiency of general-purpose language agents in self-repair, which often stems from a lack of fine-grained failure diagnosis, leading to blind context expansion and conflation of distinct error types. To overcome this, the authors propose DARC, a novel framework that prioritizes diagnosis before repair: it first analyzes failure patterns across a task family using a development set, selects appropriate repair interventions, and employs a validator to freeze the optimal success-cost strategy, thereby enforcing a causal “diagnose-then-repair” workflow. By designing recovery-oriented interfaces that integrate failure mode analysis, pruning of a shared repair library, and strategy freezing, DARC significantly improves task success rates while reducing interaction steps or retrieval overhead across diverse environments—including ALFWorld, AppWorld, and XBRL Finance—outperforming both standard foundation agents and existing general-purpose repair methods.

agent failuresdiagnostic signalsfailure modes

Hot Scholars

TG

Taylan G. Topcu

Assistant Professor of Systems Engineering & Analysis @ Virginia Tech, the Grado Department of ISE
Systems EngineeringSociotechnical SystemsDigital EngineeringModularity
QL

Qinghua Lu

Group Leader, Software Systems Research Group, CSIRO's Data61
AI EngineeringSE4AISoftware ArchitectureAI Safety
LZ

Liming Zhu

Research Director at CSIRO’s Data61 & Prof at University of New South Wales
Software ArchitectureSE4AIResponsible AIAI Safety
PL

Patricia Lago

Full Professor, S2 Group, Dept. Computer Science, Vrije Universiteit Amsterdam
Software ArchitectureSoftware EngineeringSoftware SustainabilityGreen Software
MR

Matt Roach

Swansea University
Machine LearningHCI