implement failure recovery

Designs, builds, and evaluates mechanisms and policies that detect faults and automatically restore correct operation at runtime — including automated error handling, retry and redundancy strategies, runtime/self-healing repairs, autofix-inspired automated program repair, and documented failure-recovery patterns and strategies. Implements triggers and recovery actions plus the instrumentation and logging needed to minimize time-to-recovery, maintain service throughput during faults, and enable adaptive improvement of future recovery behavior.

implementfailurerecovery

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.91
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$222K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the reliability challenges faced by modern web applications due to their inherent complexity and dynamic operating environments. The authors propose a modular self-healing framework grounded in the MAPE-K architecture, which innovatively integrates AutoFix-inspired heuristics with a learning-driven, feedback-guided recovery strategy to enable adaptive fault repair. Evaluated through fault injection experiments and iterative optimization in real-world scenarios, the system achieves an F1 score of 90.7% for fault detection and a 93.2% success rate in recovery, with an average recovery time of just 3.92 seconds. Notably, it sustains throughput at 88%–95% of baseline levels while increasing response time by only 3.1%, thereby significantly enhancing the resilience and autonomous recovery capabilities of web applications.

adaptive recoveryfault toleranceruntime failures

This work addresses the inefficiency of general-purpose language agents in self-repair, which often stems from a lack of fine-grained failure diagnosis, leading to blind context expansion and conflation of distinct error types. To overcome this, the authors propose DARC, a novel framework that prioritizes diagnosis before repair: it first analyzes failure patterns across a task family using a development set, selects appropriate repair interventions, and employs a validator to freeze the optimal success-cost strategy, thereby enforcing a causal “diagnose-then-repair” workflow. By designing recovery-oriented interfaces that integrate failure mode analysis, pruning of a shared repair library, and strategy freezing, DARC significantly improves task success rates while reducing interaction steps or retrieval overhead across diverse environments—including ALFWorld, AppWorld, and XBRL Finance—outperforming both standard foundation agents and existing general-purpose repair methods.

agent failuresdiagnostic signalsfailure modes

This work addresses a critical gap in microservice fault diagnosis: while existing methods can accurately identify root causes, they often fail to generate effective and executable recovery actions, preventing true system restoration. To bridge this gap, the authors propose R2Act, a novel framework that formally defines a recovery-oriented action space, introduces metrics for action effectiveness, and establishes an offline evaluation protocol. They also construct a benchmark dataset comprising 302 real-world Kubernetes faults, annotated with root causes and synchronized multimodal observations. Leveraging techniques such as action modeling and retrieval-augmented generation (RAG) enhanced large language models (LLMs), the study systematically evaluates the entire pipeline from diagnosis to recovery. Experimental results reveal that despite root cause localization accuracy ranging from 91.4% to 99.7%, the effectiveness of generated recovery actions remains limited at only 36.8%–60.3%, highlighting a key bottleneck in current LLM-based recovery decision-making.

diagnosis-to-action reasoningincident responselarge language models

A Survey of LLM-based Automated Program Repair: Taxonomies, Design Paradigms, and Applications

Jun 30, 2025
BY
Boyang Yang
🏛️ Jisuan Institute of Technology | Beijing JudaoYouda Network Technology Co. Ltd. | Yanshan University | China University of Mining and Technology | University of Melbourne | University of Illinois Urbana-Champaign | SnT | Nanyang Technological University

This study systematically reviews 63 LLM-based automated program repair (APR) systems published between January 2022 and June 2025, addressing three core challenges: semantic correctness verification beyond test suites, large-scale repository-level defect repair, and optimization of LLM inference cost. We propose the first comprehensive taxonomy, categorizing APR designs into four paradigms: fine-tuning, prompt engineering, pipeline-based workflows, and agent frameworks. Quantitative analysis demonstrates how retrieval augmentation and static/dynamic code analysis enhance context quality, while revealing fundamental trade-offs among cost, controllability, and scalability across paradigms. Key insights identify lightweight feedback mechanisms, repository-aware retrieval, and cost-aware planning as critical levers for advancement. The work establishes a theoretical framework and practical roadmap to enhance the reliability, scalability, and real-world applicability of LLM-APR systems.

Addressing challenges in verifying semantic correctness beyond test suitesAdvancing reliable and efficient LLM-based APR with lightweight human feedbackCategorizing LLM-based APR systems into four paradigms

Latest Papers

What's happening recently
View more

This work addresses the limited diversity in repair strategies generated by current large language models for automated program repair, which often stems from redundant execution traces and repetitive sampling. To overcome this, the authors propose CT-Repair, a novel framework that integrates static and dynamic evidence by combining Code Property Graphs (CPGs) with Temporal Execution Graphs (TEGs). CT-Repair introduces a finite state machine–guided multi-perspective agent collaboration mechanism, enabling independent generation and optimization of diverse repair strategies. Coupled with a three-stage filtering pipeline and validation feedback, the approach substantially enhances both repair diversity and accuracy. Evaluated on 854 Java bugs from Defects4J v3.0, CT-Repair successfully repairs 489, outperforming ReinFix and RepairAgent; the joint use of three perspectives yields 99 more fixes than the strongest single perspective, while execution-based filtering reduces the search space by an average of 94.85%.

Automated Program RepairExecution TracesPatch Sampling

This work addresses the unreliability of large language model (LLM)-driven agent workflows, which stems from output nondeterminism, complex node dependencies, and tool heterogeneity, and proposes FlowFixer—a novel framework that introduces symbolic reasoning into automated workflow repair. FlowFixer models execution traces symbolically to generate behavioral specifications, enabling precise fault localization and root cause identification, and dynamically synthesizes targeted repair patches. To reduce verification overhead, it incorporates a multidimensional pre-evaluation mechanism. Experimental evaluation on Dify, Coze, and n8n platforms demonstrates that FlowFixer achieves a repair success rate of 71.3%, outperforming existing methods by 11.9%–27.6%, and improves root cause analysis accuracy by 15.3%–38.8%.

agentic workflowautomatic repairfailure root cause

This work addresses the persistent reliance on manual intervention for recovering from faults in process plants that fall outside predefined monitoring logic. To enhance automation and safety, the authors propose a knowledge-guided large language model (LLM) agent framework that functions as a constrained supervisory planner. By integrating domain-specific plant knowledge, the framework generates safe recovery actions and ensures execution reliability through symbolic or simulation-based verification mechanisms. The study innovatively defines three core design dimensions for LLM agents in this context: fault recovery patterns, verification strategies, and deployment constraints. Additionally, it provides two open-source Python environments to facilitate reproduction of canonical cases and support user-defined extensions, thereby significantly advancing the automation and safety of fault recovery in industrial settings.

Fault recoveryOperator dependenceProcess plants

This work addresses the challenge of determining whether a local recovery point is semantically valid when structured tool-using agents fail mid-execution, particularly in scenarios where downstream components have already committed to outputs from upstream stages. The paper introduces DART, a runtime system that formalizes the notion of “semantic recoverability” for the first time. DART enables safe and efficient partial recovery by identifying failure instances, verifying semantic boundaries, aligning checkpoints, and selecting legitimate recovery points under dependency and effect constraints. Its modular architecture incorporates explicit acceptability checks to prevent invalidation of already-committed downstream work. Empirical evaluation across three LLM-driven tasks and the LangGraph framework demonstrates that DART successfully recovers all commitment-sensitive cases where baseline methods fail, with no unsafe rollbacks detected in a five-domain safety audit.

commitment-sensitivelocal recoveryruntime failure

Hot Scholars

ZL

Zhaoyang Liu

Tongyi Lab, Alibaba Group
LLMRecommendation
SS

Shuran Song

Stanford University
RoboticsComputer VisionMachine Learning
AS

Alberto Sonnino

Researcher, University College London
Computer SecuritySecurity EngineeringInformation SecurityPrivacy
BH

Biwei Huang

UCSD
CausalityMachine LearningComputational Science
JY

Junyan Ye

SYSU
Computer Vision and Deep Learning