build self-healing systems

Designs, builds, and analyzes systems, frameworks, pipelines, workflows, and orchestration that detect faults or degraded behavior and automatically diagnose and remediate them according to policy; includes creating automation, recovery strategies, and mechanisms that perform monitoring, root-cause analysis, rollback or corrective actions, and continuous adaptation. Implements self-healing and self‑evolving capabilities across layers of execution and deployment so systems recover or adapt autonomously to preserve correctness and availability.

buildself-healingsystems

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$199K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the frequent disruptions in data and AI pipelines caused by data anomalies, schema changes, or infrastructure failures, which existing self-healing solutions often mitigate at high cost, with vendor lock-in, or with limited adaptability for small-to-medium teams. The paper proposes an open-source, vendor-agnostic autonomous remediation architecture that integrates monitoring, metadata, historical incident logs, a policy engine, AI-driven root cause analysis, and controlled repair mechanisms to automatically detect, diagnose, remediate, and validate pipeline issues. Its key contribution lies in delivering a unified, portable end-to-end reference architecture that systematically consolidates fragmented capabilities without reliance on proprietary platforms, substantially reducing manual intervention and enhancing pipeline resilience across diverse contexts such as data engineering, MLOps, and software delivery.

AI pipelinesdata pipelinespipeline failures

This study addresses the reliability challenges faced by modern web applications due to their inherent complexity and dynamic operating environments. The authors propose a modular self-healing framework grounded in the MAPE-K architecture, which innovatively integrates AutoFix-inspired heuristics with a learning-driven, feedback-guided recovery strategy to enable adaptive fault repair. Evaluated through fault injection experiments and iterative optimization in real-world scenarios, the system achieves an F1 score of 90.7% for fault detection and a 93.2% success rate in recovery, with an average recovery time of just 3.92 seconds. Notably, it sustains throughput at 88%–95% of baseline levels while increasing response time by only 3.1%, thereby significantly enhancing the resilience and autonomous recovery capabilities of web applications.

adaptive recoveryfault toleranceruntime failures

To address the delayed fault response and high manual remediation costs in modern software systems, this paper proposes an AI-driven self-healing software architecture inspired by biological autorepair mechanisms. Methodologically, it formalizes the human injury-sensing–diagnosis–repair paradigm into a novel “Perceive–Cognize–Execute” three-tier closed-loop framework; integrates system observability signals, multimodal log analysis, static code inspection, and large language model (LLM)-based generation to enable automated fault localization, patch/test synthesis, and lightweight agent-driven repair execution. Contributions include the first bio-inspired self-healing architecture paradigm and a scalable, collaborative self-healing technology stack. Experimental evaluation demonstrates a significant reduction in mean time to recovery (MTTR), a 3.2× improvement in debugging efficiency, and a 76% decrease in critical service outages, empirically validating sustained self-healing capability.

Combining log analysis and AI to automate debugging and patchingDeveloping AI-driven self-healing software systems for failure recoveryMimicking biological healing to reduce downtime and enhance resilience

To address insufficient resilience of complex systems under heterogeneous hardware environments, this paper proposes a fault-adaptive software deployment and redundancy configuration optimization method. We construct a system-level resilience state-space model and introduce a novel equivalence relation to enable quotient-space-based state-space reduction, significantly compressing the state space. Subsequently, we integrate formal model checking with strategy synthesis to automatically derive both an initial deployment configuration and dynamic reconfiguration policies that satisfy multi-level resilience requirements. Our key contributions are: (i) a new equivalence relation enabling efficient, semantics-preserving state-space reduction; and (ii) end-to-end automated synthesis of fault-response and recovery strategies. Experimental evaluation on an autonomous driving system model demonstrates that our approach substantially improves fault recovery latency and system availability, while supporting real-time resilience assurance.

Automated framework for resilient complex systems under failuresGenerates resilient initial configurations and reconfiguration policiesOptimized adaptive distribution and replication of software components

This work addresses the lack of closed-loop control in traditional software development lifecycles, which often fails to simultaneously ensure security, auditability, and highly reliable automation. The authors propose a deterministic autonomous control framework that models the lifecycle as a seven-stage automated pipeline, integrating Jira-based task orchestration, structured context, resource constraints, and human-review gating mechanisms to establish a secure closed loop. Key innovations include a state-contract-based collision locking mechanism, a degradation protocol for fallback operation, and a traceable control architecture. Implemented with 12,661 lines of Python code and 6,907 lines of versioned prompt specifications—including 101 exception handlers and 12 centralized locks—the system achieved a 100% success rate (95% CI [97.6%, 100%]) across 152 initial runs, producing over 795 artifacts. All 51 issues identified through adversarial review were fully resolved, with 60% of security tickets autonomously completed.

Autonomous Software DevelopmentBacklog OrchestrationClosed-Loop Control

Latest Papers

What's happening recently
View more

This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).

Agentic SystemsMonitoringStructural Defects

This work addresses the unreliability of large language model (LLM)-driven agent workflows, which stems from output nondeterminism, complex node dependencies, and tool heterogeneity, and proposes FlowFixer—a novel framework that introduces symbolic reasoning into automated workflow repair. FlowFixer models execution traces symbolically to generate behavioral specifications, enabling precise fault localization and root cause identification, and dynamically synthesizes targeted repair patches. To reduce verification overhead, it incorporates a multidimensional pre-evaluation mechanism. Experimental evaluation on Dify, Coze, and n8n platforms demonstrates that FlowFixer achieves a repair success rate of 71.3%, outperforming existing methods by 11.9%–27.6%, and improves root cause analysis accuracy by 15.3%–38.8%.

agentic workflowautomatic repairfailure root cause

This work addresses the inefficiency of general-purpose language agents in self-repair, which often stems from a lack of fine-grained failure diagnosis, leading to blind context expansion and conflation of distinct error types. To overcome this, the authors propose DARC, a novel framework that prioritizes diagnosis before repair: it first analyzes failure patterns across a task family using a development set, selects appropriate repair interventions, and employs a validator to freeze the optimal success-cost strategy, thereby enforcing a causal “diagnose-then-repair” workflow. By designing recovery-oriented interfaces that integrate failure mode analysis, pruning of a shared repair library, and strategy freezing, DARC significantly improves task success rates while reducing interaction steps or retrieval overhead across diverse environments—including ALFWorld, AppWorld, and XBRL Finance—outperforming both standard foundation agents and existing general-purpose repair methods.

agent failuresdiagnostic signalsfailure modes

This work addresses the persistent reliance on manual intervention for recovering from faults in process plants that fall outside predefined monitoring logic. To enhance automation and safety, the authors propose a knowledge-guided large language model (LLM) agent framework that functions as a constrained supervisory planner. By integrating domain-specific plant knowledge, the framework generates safe recovery actions and ensures execution reliability through symbolic or simulation-based verification mechanisms. The study innovatively defines three core design dimensions for LLM agents in this context: fault recovery patterns, verification strategies, and deployment constraints. Additionally, it provides two open-source Python environments to facilitate reproduction of canonical cases and support user-defined extensions, thereby significantly advancing the automation and safety of fault recovery in industrial settings.

Fault recoveryOperator dependenceProcess plants

Hot Scholars

CL

Cong Lu

Google DeepMind
Reinforcement LearningOpen-EndednessGenerative ModelingDeep Learning
AA

Andrew Adamatzky

Professor in Unconventional Computing, UWE, Bristol
computer scienceunconventional computing novel materials cellular automata applied mathematicsrobotics
LB

Lei Bai

Shanghai AI Laboratory
Foundation ModelScience IntelligenceMulti-Agent SystemAutonomous Discovery
FK

Fang Kong

Southern University of Science and Technology, Assistant Professor
multi-armed banditsonline learningreinforcement learning
GW

Guancheng Wan

Computer Science, UCLA
AI AgentAI4ScienceLarge Language ModelTrustworthy AI