Score
Designs, builds, and evaluates mechanisms and policies that detect faults and automatically restore correct operation at runtime — including automated error handling, retry and redundancy strategies, runtime/self-healing repairs, autofix-inspired automated program repair, and documented failure-recovery patterns and strategies. Implements triggers and recovery actions plus the instrumentation and logging needed to minimize time-to-recovery, maintain service throughput during faults, and enable adaptive improvement of future recovery behavior.
Microservice systems commonly exhibit resilience deficiencies—including localized fault propagation, cascading timeouts, and inconsistent recovery behaviors—yet existing research remains largely descriptive, lacking systematic evidence synthesis and quantitative evaluation. To address this gap, we conduct the first PRISMA-guided systematic literature review (SLR) of 26 high-quality empirical studies published between 2014 and 2025. Our analysis identifies nine core recovery patterns and introduces three novel, empirically grounded artifacts: (1) a reproducible Recovery Pattern Taxonomy; (2) a standardized Resilience Evaluation Score for quantitative assessment; and (3) a constraint-aware decision matrix that explicitly trades off latency, consistency, and cost. Collectively, these contributions establish a structured, quantifiable, and reproducible empirical foundation for resilience-aware microservice design and engineering.
This study addresses the reliability challenges faced by modern web applications due to their inherent complexity and dynamic operating environments. The authors propose a modular self-healing framework grounded in the MAPE-K architecture, which innovatively integrates AutoFix-inspired heuristics with a learning-driven, feedback-guided recovery strategy to enable adaptive fault repair. Evaluated through fault injection experiments and iterative optimization in real-world scenarios, the system achieves an F1 score of 90.7% for fault detection and a 93.2% success rate in recovery, with an average recovery time of just 3.92 seconds. Notably, it sustains throughput at 88%–95% of baseline levels while increasing response time by only 3.1%, thereby significantly enhancing the resilience and autonomous recovery capabilities of web applications.
This work addresses the inefficiency of general-purpose language agents in self-repair, which often stems from a lack of fine-grained failure diagnosis, leading to blind context expansion and conflation of distinct error types. To overcome this, the authors propose DARC, a novel framework that prioritizes diagnosis before repair: it first analyzes failure patterns across a task family using a development set, selects appropriate repair interventions, and employs a validator to freeze the optimal success-cost strategy, thereby enforcing a causal “diagnose-then-repair” workflow. By designing recovery-oriented interfaces that integrate failure mode analysis, pruning of a shared repair library, and strategy freezing, DARC significantly improves task success rates while reducing interaction steps or retrieval overhead across diverse environments—including ALFWorld, AppWorld, and XBRL Finance—outperforming both standard foundation agents and existing general-purpose repair methods.
This work addresses a critical gap in microservice fault diagnosis: while existing methods can accurately identify root causes, they often fail to generate effective and executable recovery actions, preventing true system restoration. To bridge this gap, the authors propose R2Act, a novel framework that formally defines a recovery-oriented action space, introduces metrics for action effectiveness, and establishes an offline evaluation protocol. They also construct a benchmark dataset comprising 302 real-world Kubernetes faults, annotated with root causes and synchronized multimodal observations. Leveraging techniques such as action modeling and retrieval-augmented generation (RAG) enhanced large language models (LLMs), the study systematically evaluates the entire pipeline from diagnosis to recovery. Experimental results reveal that despite root cause localization accuracy ranging from 91.4% to 99.7%, the effectiveness of generated recovery actions remains limited at only 36.8%–60.3%, highlighting a key bottleneck in current LLM-based recovery decision-making.
This study systematically reviews 63 LLM-based automated program repair (APR) systems published between January 2022 and June 2025, addressing three core challenges: semantic correctness verification beyond test suites, large-scale repository-level defect repair, and optimization of LLM inference cost. We propose the first comprehensive taxonomy, categorizing APR designs into four paradigms: fine-tuning, prompt engineering, pipeline-based workflows, and agent frameworks. Quantitative analysis demonstrates how retrieval augmentation and static/dynamic code analysis enhance context quality, while revealing fundamental trade-offs among cost, controllability, and scalability across paradigms. Key insights identify lightweight feedback mechanisms, repository-aware retrieval, and cost-aware planning as critical levers for advancement. The work establishes a theoretical framework and practical roadmap to enhance the reliability, scalability, and real-world applicability of LLM-APR systems.
This work addresses the limited diversity in repair strategies generated by current large language models for automated program repair, which often stems from redundant execution traces and repetitive sampling. To overcome this, the authors propose CT-Repair, a novel framework that integrates static and dynamic evidence by combining Code Property Graphs (CPGs) with Temporal Execution Graphs (TEGs). CT-Repair introduces a finite state machine–guided multi-perspective agent collaboration mechanism, enabling independent generation and optimization of diverse repair strategies. Coupled with a three-stage filtering pipeline and validation feedback, the approach substantially enhances both repair diversity and accuracy. Evaluated on 854 Java bugs from Defects4J v3.0, CT-Repair successfully repairs 489, outperforming ReinFix and RepairAgent; the joint use of three perspectives yields 99 more fixes than the strongest single perspective, while execution-based filtering reduces the search space by an average of 94.85%.
This work addresses the unreliability of large language model (LLM)-driven agent workflows, which stems from output nondeterminism, complex node dependencies, and tool heterogeneity, and proposes FlowFixer—a novel framework that introduces symbolic reasoning into automated workflow repair. FlowFixer models execution traces symbolically to generate behavioral specifications, enabling precise fault localization and root cause identification, and dynamically synthesizes targeted repair patches. To reduce verification overhead, it incorporates a multidimensional pre-evaluation mechanism. Experimental evaluation on Dify, Coze, and n8n platforms demonstrates that FlowFixer achieves a repair success rate of 71.3%, outperforming existing methods by 11.9%–27.6%, and improves root cause analysis accuracy by 15.3%–38.8%.
This work addresses the persistent reliance on manual intervention for recovering from faults in process plants that fall outside predefined monitoring logic. To enhance automation and safety, the authors propose a knowledge-guided large language model (LLM) agent framework that functions as a constrained supervisory planner. By integrating domain-specific plant knowledge, the framework generates safe recovery actions and ensures execution reliability through symbolic or simulation-based verification mechanisms. The study innovatively defines three core design dimensions for LLM agents in this context: fault recovery patterns, verification strategies, and deployment constraints. Additionally, it provides two open-source Python environments to facilitate reproduction of canonical cases and support user-defined extensions, thereby significantly advancing the automation and safety of fault recovery in industrial settings.
This work addresses the challenge of determining whether a local recovery point is semantically valid when structured tool-using agents fail mid-execution, particularly in scenarios where downstream components have already committed to outputs from upstream stages. The paper introduces DART, a runtime system that formalizes the notion of “semantic recoverability” for the first time. DART enables safe and efficient partial recovery by identifying failure instances, verifying semantic boundaries, aligning checkpoints, and selecting legitimate recovery points under dependency and effect constraints. Its modular architecture incorporates explicit acceptability checks to prevent invalidation of already-committed downstream work. Empirical evaluation across three LLM-driven tasks and the LangGraph framework demonstrates that DART successfully recovers all commitment-sensitive cases where baseline methods fail, with no unsafe rollbacks detected in a five-domain safety audit.