Score
Designs, builds, or analyzes systems, architectures, materials, processes, or organizational practices to maintain required functionality under stress, to degrade gracefully, and to recover quickly after disturbances. Work includes specifying failure modes and resilience metrics, implementing redundancy and recovery mechanisms, and evaluating tradeoffs among robustness, adaptability, and recovery time.
In cloud environments, selecting optimal data protection strategies for business continuity and disaster recovery remains challenging due to the lack of quantitative foundations for evaluating reliability and aligning with organizational Recovery Time Objectives (RTOs) and operational requirements. Method: This paper proposes an integrated assessment framework that synergistically combines system dynamics modeling and simulation-based optimization. It quantitatively evaluates key performance indicators—including recovery timeliness, data integrity, and system robustness—across public and hybrid cloud scenarios by simulating mainstream recovery mechanisms. Contribution/Results: The framework innovatively applies system dynamics to model time-varying dependencies during recovery processes and establishes interpretable, traceable mappings between policy parameters, technical metrics, and business objectives. Empirical validation demonstrates its reproducibility and practical utility, providing cloud-native organizations with a quantifiable, verifiable, and actionable decision-support methodology for data protection strategy selection.
To address insufficient resilience of complex systems under heterogeneous hardware environments, this paper proposes a fault-adaptive software deployment and redundancy configuration optimization method. We construct a system-level resilience state-space model and introduce a novel equivalence relation to enable quotient-space-based state-space reduction, significantly compressing the state space. Subsequently, we integrate formal model checking with strategy synthesis to automatically derive both an initial deployment configuration and dynamic reconfiguration policies that satisfy multi-level resilience requirements. Our key contributions are: (i) a new equivalence relation enabling efficient, semantics-preserving state-space reduction; and (ii) end-to-end automated synthesis of fault-response and recovery strategies. Experimental evaluation on an autonomous driving system model demonstrates that our approach substantially improves fault recovery latency and system availability, while supporting real-time resilience assurance.
Chaos engineering lacks a systematic, comprehensive review in the literature. Method: This paper conducts the first multi-source literature review (MLR), systematically analyzing 96 academic and gray literature sources published between 2016 and 2024—including 88 core publications from 2019 to 2024. It synthesizes findings via thematic clustering and qualitative coding. Contribution/Results: The study establishes the first consensus definition of chaos engineering, proposes a four-layer capability model and a five-dimensional component taxonomy, and performs a cross-tool evaluation of 12 mainstream chaos engineering tools. It identifies six open research challenges and clarifies practice drivers, tool characteristics, and research evolution trends. The results provide a foundational theoretical framework, methodological benchmark, and roadmap for future work—bridging critical knowledge gaps between academia and industry.
Current cyber-physical systems (CPS) in vehicular environments lack quantitative, experimentally grounded methods for assessing network resilience. Method: This study constructs an experimental testbed replicating real-world truck operational conditions and conducts multiple rounds of malware injection attacks, simultaneously collecting network- and physical-layer data on resistance and recovery behaviors. Contribution/Results: We introduce the novel concept of “bonware” to holistically characterize both cybersecurity defense capability and physical resilience, formalized via an analytically tractable mathematical model. We further define and extract experimentally identifiable, quantitative resilience metrics—termed elastic features—for the first time. Sensitivity analysis confirms these metrics exhibit significant discriminability with respect to attack intensity, defensive strategies, and physical redundancy. This work bridges a critical gap by advancing vehicular CPS resilience from qualitative description to quantifiable, comparable, and optimizable measurement.
Energy-sector industrial control systems (ICS) exhibit insufficient security resilience and overreliance on reactive, post-incident remediation. Method: This paper proposes a layered, implementable Security-by-Design (SbD) framework and a deployable set of security requirements tailored to critical infrastructure. Integrating systems engineering, ICS-specific security architecture, organizational behavior principles, and continuous monitoring, the approach spans the entire lifecycle—design, development, deployment, and operations—while ensuring alignment with IEC 62443 and NIST SP 800-82. Contribution/Results: It represents the first systematic, end-to-end operationalization of SbD in energy ICS contexts, enabling a paradigm shift from passive incident response to inherent, “native immunity.” The resulting scalable, auditable, and standards-coordinated SbD implementation guide supports the development of high-assurance, resilient, and sustainably evolvable cybersecurity ecosystems.
This study addresses occupational burnout among Security Operations Center (SOC) practitioners, often stemming from misalignment between job demands and individual capabilities. Drawing on flow theory, the authors conduct an inductive content analysis of 106 global SOC job postings to systematically map the prevalence of certifications (e.g., CISSP), technical skills (e.g., Python, Splunk), and soft skills—particularly communication skills, mentioned in 50.9% of listings. The research reveals, for the first time, a structured pattern in the skill and certification requirements of SOC roles. These findings provide empirical grounding for achieving challenge–skill balance, refining recruitment practices, and guiding professional development. Furthermore, the study advances the discourse on flow-aligned person–job fit and sets the stage for future investigations into the impact of artificial intelligence on SOC workforce dynamics.
This study addresses the inadequacy of current IT compliance–oriented cybersecurity policies in safeguarding the physical safety of cyber-physical systems, as digital failures often precipitate real-world harm. By coding 292 critical infrastructure policies (2000–2025) and aligning them with the NIST SP 800-160 Vol. 2 resilience lifecycle, the research reveals a significant misalignment between prevailing policy approaches—overreliant on IT control catalogs during resistance and recovery phases—and actual physical risks. The work proposes a modernized “duty of reasonable care” standard centered on hazard-specific traceability, structured assurance cases, and cyber resilience engineering. It identifies three critical disconnects: misaligned delegation of standards, reduction of recovery mechanisms to mere incident reporting, and uneven sectoral adaptability. The study further outlines a viable pathway for federal policy that integrates engineering implementation with targeted incentives.
This study addresses the vulnerability of complex systems to self-induced collapse under shocks, a challenge inadequately met by conventional approaches that struggle to balance robustness and adaptability. The work distinguishes two dynamic regimes—phase-separated systems and highly volatile, interwoven systems—and conceptualizes resilience as an emergent property arising from multi-agent interactions. Rather than advocating mere restoration to prior states, it proposes systemic transformation to enhance recovery capacity. Methodologically, the research integrates data-driven multi-agent modeling, knowledge graphs, and artificial intelligence tools. Large-scale simulations reveal that optimizing for peak performance often undermines resilience, whereas second-order interventions leveraging positive feedback mechanisms can effectively reconfigure system architecture, thereby substantially strengthening overall resilience.
This study addresses the critical challenge of assessing project vulnerability due to personnel attrition, a risk that existing methods often underestimate or oversimplify. To this end, the work introduces network robustness theory into project resilience analysis for the first time, proposing a task–personnel bipartite network model that integrates complex network modeling, node importance evaluation, and connectivity metrics to capture both task dependencies and personnel allocation structures. By explicitly accounting for structural fragmentation and avoiding overly optimistic assumptions inherent in conventional approaches, the proposed framework delivers more consistent and accurate quantification of project resilience across diverse scenarios. It effectively identifies mission-critical personnel and enables proactive prediction of potential disruption risks arising from workforce instability.
This study addresses the critical gap in existing Software Engineering for AI (SE4AI) methodologies regarding the irreversible physical consequences of failures in AI-driven building operations. To overcome this limitation, we propose a novel SE4AI paradigm specifically tailored for systems with physical implications. Through interdisciplinary research and case analysis, this work establishes a foundational software engineering framework and best practice guidelines applicable to physical systems. By rectifying the oversight of traditional methods concerning energy consumption and equipment degradation, this research provides essential theoretical support and engineering standards for the safe and reliable deployment of AI in building automation. Ultimately, these contributions advance the innovation and development of software engineering methodologies within this domain, ensuring robust integration of intelligent technologies in critical infrastructure.