Score
Designs, builds, and evaluates mechanisms and policies that detect faults and automatically restore correct operation at runtime — including automated error handling, retry and redundancy strategies, runtime/self-healing repairs, autofix-inspired automated program repair, and documented failure-recovery patterns and strategies. Implements triggers and recovery actions plus the instrumentation and logging needed to minimize time-to-recovery, maintain service throughput during faults, and enable adaptive improvement of future recovery behavior.
Microservice systems commonly exhibit resilience deficiencies—including localized fault propagation, cascading timeouts, and inconsistent recovery behaviors—yet existing research remains largely descriptive, lacking systematic evidence synthesis and quantitative evaluation. To address this gap, we conduct the first PRISMA-guided systematic literature review (SLR) of 26 high-quality empirical studies published between 2014 and 2025. Our analysis identifies nine core recovery patterns and introduces three novel, empirically grounded artifacts: (1) a reproducible Recovery Pattern Taxonomy; (2) a standardized Resilience Evaluation Score for quantitative assessment; and (3) a constraint-aware decision matrix that explicitly trades off latency, consistency, and cost. Collectively, these contributions establish a structured, quantifiable, and reproducible empirical foundation for resilience-aware microservice design and engineering.
This study addresses the reliability challenges faced by modern web applications due to their inherent complexity and dynamic operating environments. The authors propose a modular self-healing framework grounded in the MAPE-K architecture, which innovatively integrates AutoFix-inspired heuristics with a learning-driven, feedback-guided recovery strategy to enable adaptive fault repair. Evaluated through fault injection experiments and iterative optimization in real-world scenarios, the system achieves an F1 score of 90.7% for fault detection and a 93.2% success rate in recovery, with an average recovery time of just 3.92 seconds. Notably, it sustains throughput at 88%–95% of baseline levels while increasing response time by only 3.1%, thereby significantly enhancing the resilience and autonomous recovery capabilities of web applications.
This work addresses the inefficiency of general-purpose language agents in self-repair, which often stems from a lack of fine-grained failure diagnosis, leading to blind context expansion and conflation of distinct error types. To overcome this, the authors propose DARC, a novel framework that prioritizes diagnosis before repair: it first analyzes failure patterns across a task family using a development set, selects appropriate repair interventions, and employs a validator to freeze the optimal success-cost strategy, thereby enforcing a causal “diagnose-then-repair” workflow. By designing recovery-oriented interfaces that integrate failure mode analysis, pruning of a shared repair library, and strategy freezing, DARC significantly improves task success rates while reducing interaction steps or retrieval overhead across diverse environments—including ALFWorld, AppWorld, and XBRL Finance—outperforming both standard foundation agents and existing general-purpose repair methods.
本文针对DevOps部署中的自动故障恢复难题,提出了一种结合基于规则和机器学习的混合框架,通过实时监控、识别、分类并恢复故障,有效提高了系统的恢复效率与运行稳定性。
论文提出一种框架明确指定基于大语言模型的程序修复实验设置,通过分析Defects4J和SWE-bench上的系统,解决了实验设置不透明导致的结果可比性问题。
This work addresses a critical gap in microservice fault diagnosis: while existing methods can accurately identify root causes, they often fail to generate effective and executable recovery actions, preventing true system restoration. To bridge this gap, the authors propose R2Act, a novel framework that formally defines a recovery-oriented action space, introduces metrics for action effectiveness, and establishes an offline evaluation protocol. They also construct a benchmark dataset comprising 302 real-world Kubernetes faults, annotated with root causes and synchronized multimodal observations. Leveraging techniques such as action modeling and retrieval-augmented generation (RAG) enhanced large language models (LLMs), the study systematically evaluates the entire pipeline from diagnosis to recovery. Experimental results reveal that despite root cause localization accuracy ranging from 91.4% to 99.7%, the effectiveness of generated recovery actions remains limited at only 36.8%–60.3%, highlighting a key bottleneck in current LLM-based recovery decision-making.
This work addresses the limited diversity in repair strategies generated by current large language models for automated program repair, which often stems from redundant execution traces and repetitive sampling. To overcome this, the authors propose CT-Repair, a novel framework that integrates static and dynamic evidence by combining Code Property Graphs (CPGs) with Temporal Execution Graphs (TEGs). CT-Repair introduces a finite state machine–guided multi-perspective agent collaboration mechanism, enabling independent generation and optimization of diverse repair strategies. Coupled with a three-stage filtering pipeline and validation feedback, the approach substantially enhances both repair diversity and accuracy. Evaluated on 854 Java bugs from Defects4J v3.0, CT-Repair successfully repairs 489, outperforming ReinFix and RepairAgent; the joint use of three perspectives yields 99 more fixes than the strongest single perspective, while execution-based filtering reduces the search space by an average of 94.85%.
This work addresses the unreliability of large language model (LLM)-driven agent workflows, which stems from output nondeterminism, complex node dependencies, and tool heterogeneity, and proposes FlowFixer—a novel framework that introduces symbolic reasoning into automated workflow repair. FlowFixer models execution traces symbolically to generate behavioral specifications, enabling precise fault localization and root cause identification, and dynamically synthesizes targeted repair patches. To reduce verification overhead, it incorporates a multidimensional pre-evaluation mechanism. Experimental evaluation on Dify, Coze, and n8n platforms demonstrates that FlowFixer achieves a repair success rate of 71.3%, outperforming existing methods by 11.9%–27.6%, and improves root cause analysis accuracy by 15.3%–38.8%.
研究针对LLM代理软件中审批与执行间不一致问题,提出通过重构稳定授权、衍生义务及APAS-Finder工具实现修复分析,确保授权准确无误。
ORCA通过基于观测性的方法,将微服务故障诊断与程序修复相结合,利用故障特征定位问题代码并生成修复补丁,有效解决了从诊断到修复的转换难题。
This work addresses the persistent reliance on manual intervention for recovering from faults in process plants that fall outside predefined monitoring logic. To enhance automation and safety, the authors propose a knowledge-guided large language model (LLM) agent framework that functions as a constrained supervisory planner. By integrating domain-specific plant knowledge, the framework generates safe recovery actions and ensures execution reliability through symbolic or simulation-based verification mechanisms. The study innovatively defines three core design dimensions for LLM agents in this context: fault recovery patterns, verification strategies, and deployment constraints. Additionally, it provides two open-source Python environments to facilitate reproduction of canonical cases and support user-defined extensions, thereby significantly advancing the automation and safety of fault recovery in industrial settings.