Score
Designs, implements, and operates backup and recovery systems and processes, including automated and on‑demand backups, retention and lifecycle policies, encryption, storage placement (including offsite/secondary stores), and replication mechanisms. Builds and verifies recovery and failover procedures, scripts, and tests to restore data and systems to defined recovery point and time objectives while ensuring data integrity and consistency.
In cloud environments, selecting optimal data protection strategies for business continuity and disaster recovery remains challenging due to the lack of quantitative foundations for evaluating reliability and aligning with organizational Recovery Time Objectives (RTOs) and operational requirements. Method: This paper proposes an integrated assessment framework that synergistically combines system dynamics modeling and simulation-based optimization. It quantitatively evaluates key performance indicators—including recovery timeliness, data integrity, and system robustness—across public and hybrid cloud scenarios by simulating mainstream recovery mechanisms. Contribution/Results: The framework innovatively applies system dynamics to model time-varying dependencies during recovery processes and establishes interpretable, traceable mappings between policy parameters, technical metrics, and business objectives. Empirical validation demonstrates its reproducibility and practical utility, providing cloud-native organizations with a quantifiable, verifiable, and actionable decision-support methodology for data protection strategy selection.
Modern data storage systems suffer from latent cross-layer faults due to tight hardware–software coupling across multiple abstraction layers, often leading to silent data corruption or unrecoverable data loss. To address this, we propose the first cross-layer fault-tolerance analysis framework targeting heterogeneous storage stacks—including SSDs, persistent memory, local file systems, and distributed storage. Our approach combines architectural modeling of the full stack, systematic injection of representative defects, and precise tracking of fault propagation across hardware–firmware–software boundaries to expose error propagation paths and consistency violation mechanisms. Through empirical evaluation across widely deployed systems, we identify critical vulnerabilities impacting data integrity and quantify coverage gaps in existing fault-tolerance techniques. The framework provides a scalable, principled methodology for analyzing cross-layer resilience and establishes concrete, actionable directions for designing next-generation highly reliable storage systems.
研究了AI代理系统中检查点和回滚(C/R)的安全性问题,通过建立执行模型识别出五种基本故障模式,并展示了它们对安全的影响。
In cloud computing, untrusted storage interfaces are vulnerable to rollback and replay attacks, compromising the integrity of confidential application decisions. Existing hardware-enforced state continuity mechanisms treat all rollbacks as malicious, thus precluding legitimate rollbacks required for fault recovery. This paper introduces Rebound, the first framework enabling fine-grained differentiation between malicious and legitimate rollbacks in confidential cloud environments. Rebound features a policy-authorized reference monitor—hardware-protected via trusted execution—supporting atomic state updates, controlled rollback, and tamper-evident audit logging. Leveraging formal verification, policy-driven access control, and end-to-end logging, we evaluate Rebound in GitLab CI, demonstrating efficient and secure version management for binaries, configurations, and data. The system incurs low overhead while providing strong security guarantees against unauthorized or inconsistent state transitions.
In digital forensics, the atomicity and integrity of storage snapshots lack rigorous definitions that jointly guarantee both instantaneousness and causal ordering—undermining evidentiary admissibility in legal proceedings. To address this, we propose a novel atomicity definition grounded in causal consistency, overcoming the limitation of conventional time-based atomicity models. We further rectify conceptual flaws in existing integrity definitions and introduce a revised, theoretically sound yet engineering-practical integrity criterion—explicitly supporting copy-on-write (CoW) implementations. Our approach integrates causal modeling, formal snapshot semantics, CoW mechanism analysis, and formalization of forensic quality criteria, yielding a verifiable snapshot semantic framework. This work establishes the first theoretical foundation for forensic tool design that unifies causal ordering with instantaneous state capture, thereby significantly enhancing the forensic validity and judicial admissibility of live data acquisition.
This study addresses the critical challenge that restoring IT backups alone is insufficient to resume production after ransomware attacks on manufacturing systems, due to deep interdependencies among IT, operational technology (OT), physical processes, identity management, and supply chains. The work reframes recovery as a problem of interdependent continuity in critical infrastructure and introduces, for the first time, the concept of “Minimum Viable Factory Recovery” (MVF Recovery), shifting the objective from full-system restoration to capability-oriented minimal trusted operations. Drawing on a PRISMA-guided multi-source systematic review integrating academic literature, standards, government guidelines, and real-world incidents, the study identifies nine failure modes in recovery efforts and proposes a capability-centered recovery framework. It further establishes an evidence-driven recovery lifecycle model and outlines directions for benchmarking, offering actionable recovery targets for critical manufacturing infrastructure.
This study addresses the absence of governance frameworks for digital legacies after a user’s death and the coordination difficulties faced by bereaved survivors. Through a multidimensional content analysis of 800 Reddit posts, it investigates sociotechnical collaboration and conflict surrounding post-mortem digital privacy and security. The findings reveal the critical role of devices as access gateways and identify data loss as a primary harm. Building on these insights, the authors construct a post-mortem digital governance framework organized along dimensions such as assets and actors. This framework establishes a foundation for mechanism design to support survivor coordination, thereby bridging significant theoretical and practical gaps in the field of digital legacy governance.
研究解决了AI代理在执行多步骤任务时中断后如何安全恢复的问题,通过引入可恢复性作为系统原语,明确恢复点和恢复动作,并通过验证和独立检查确保安全。
This work addresses the challenge of achieving exactly-once semantics in distributed systems recovering from crashes under a shared-nothing architecture. By formally modeling the information-theoretic boundaries inherent in crash recovery, it demonstrates that relying solely on locally persisted state inevitably leads to either duplicate or missed deliveries. The paper presents the first rigorous, machine-checkable definition of exactly-once semantics together with precise conditions for its realizability. Leveraging a formal framework built in Isabelle/HOL that integrates information-theoretic limits, crash timing, receiver-side fencing, and evidence validity analysis, the study proves that conventional dual-write protocols suffer from fundamental uncertainty. It then introduces a provably safe recovery mechanism and quantifies how memory constraints for deduplication and source log truncation impact the guaranteed validity window.
Existing workflow persistence frameworks lack precise, machine-verifiable recovery semantic contracts, often resulting in inconsistent post-crash behaviors or violations of their own guarantees. This work proposes RESUME CONTRACT, which formally specifies six core recovery properties and employs TLA+ modeling alongside Verus verification to establish their independence and correctness. Leveraging a deterministic testing framework, the authors empirically evaluate prominent systems across a state space of 7.4 million configurations, uncovering semantic flaws in widely used frameworks such as LangGraph and CrewAI. Guided by these findings, they develop REMIT, a reference implementation that effectively addresses critical issues including fork semantics, recovery validity, and cross-process duplicate consumption, and has already been successfully deployed.