Score
Designs and implements tests, simulations, and evaluation frameworks that determine whether specified recovery or remediation plans and their actions will restore system health and satisfy operational constraints. This includes defining validation metrics, replaying or simulating actions in controlled or live-like environments, checking action admissibility against policies and constraints, and analyzing outcomes to accept, reject, or refine recovery plans.
In cloud environments, selecting optimal data protection strategies for business continuity and disaster recovery remains challenging due to the lack of quantitative foundations for evaluating reliability and aligning with organizational Recovery Time Objectives (RTOs) and operational requirements. Method: This paper proposes an integrated assessment framework that synergistically combines system dynamics modeling and simulation-based optimization. It quantitatively evaluates key performance indicators—including recovery timeliness, data integrity, and system robustness—across public and hybrid cloud scenarios by simulating mainstream recovery mechanisms. Contribution/Results: The framework innovatively applies system dynamics to model time-varying dependencies during recovery processes and establishes interpretable, traceable mappings between policy parameters, technical metrics, and business objectives. Empirical validation demonstrates its reproducibility and practical utility, providing cloud-native organizations with a quantifiable, verifiable, and actionable decision-support methodology for data protection strategy selection.
This work addresses the challenge of costly erroneous repairs in existing automated remediation systems, which often lack the ability to assess intervention necessity and thus rely on manual approval for safety. The authors formulate safe repair as an intervention decision problem under risk constraints and introduce a three-dimensional risk decomposition framework encompassing impact scope, reversibility, and epistemic uncertainty. They further design a context-adaptive human-in-the-loop gating strategy that enables interpretable, workload-aware safety interventions. Built upon constrained Markov decision processes (CMDPs), offline policy learning, Chaos Mesh fault injection, and the RCAEval classification framework, the proposed approach reduces erroneous repair rates by 39% and improves repair success rates by 2.5 percentage points on the Train Ticket benchmark, while decreasing on-call escalation burden by 17% compared to fixed-threshold baselines.
This work addresses the persistent reliance on manual intervention for recovering from faults in process plants that fall outside predefined monitoring logic. To enhance automation and safety, the authors propose a knowledge-guided large language model (LLM) agent framework that functions as a constrained supervisory planner. By integrating domain-specific plant knowledge, the framework generates safe recovery actions and ensures execution reliability through symbolic or simulation-based verification mechanisms. The study innovatively defines three core design dimensions for LLM agents in this context: fault recovery patterns, verification strategies, and deployment constraints. Additionally, it provides two open-source Python environments to facilitate reproduction of canonical cases and support user-defined extensions, thereby significantly advancing the automation and safety of fault recovery in industrial settings.
This work addresses a critical gap in microservice fault diagnosis: while existing methods can accurately identify root causes, they often fail to generate effective and executable recovery actions, preventing true system restoration. To bridge this gap, the authors propose R2Act, a novel framework that formally defines a recovery-oriented action space, introduces metrics for action effectiveness, and establishes an offline evaluation protocol. They also construct a benchmark dataset comprising 302 real-world Kubernetes faults, annotated with root causes and synchronized multimodal observations. Leveraging techniques such as action modeling and retrieval-augmented generation (RAG) enhanced large language models (LLMs), the study systematically evaluates the entire pipeline from diagnosis to recovery. Experimental results reveal that despite root cause localization accuracy ranging from 91.4% to 99.7%, the effectiveness of generated recovery actions remains limited at only 36.8%–60.3%, highlighting a key bottleneck in current LLM-based recovery decision-making.
This work proposes an applicability-aware surrogate model for black-box security scoring engines, enabling accurate prediction of how remediation actions affect an organization’s security score without revealing the engine’s internal logic. The approach explicitly models the applicability of individual security checks and integrates sensitivity analysis with a reliability assessment layer to identify scenarios where predictions may be unstable. Evaluated on a real-world dataset comprising 5,188 organizational configurations, the proposed model significantly outperforms baseline methods relying on simplistic feature representations. It not only enhances the accuracy of security score predictions but also effectively flags remediation impact estimates that warrant cautious interpretation due to potential unreliability.
This work addresses the inefficiency of general-purpose language agents in self-repair, which often stems from a lack of fine-grained failure diagnosis, leading to blind context expansion and conflation of distinct error types. To overcome this, the authors propose DARC, a novel framework that prioritizes diagnosis before repair: it first analyzes failure patterns across a task family using a development set, selects appropriate repair interventions, and employs a validator to freeze the optimal success-cost strategy, thereby enforcing a causal “diagnose-then-repair” workflow. By designing recovery-oriented interfaces that integrate failure mode analysis, pruning of a shared repair library, and strategy freezing, DARC significantly improves task success rates while reducing interaction steps or retrieval overhead across diverse environments—including ALFWorld, AppWorld, and XBRL Finance—outperforming both standard foundation agents and existing general-purpose repair methods.