An Intelligent Fault Self-Healing Mechanism for Cloud AI Systems via Integration of Large Language Models and Deep Reinforcement Learning

📅 2025-06-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the growing challenge of fault detection and adaptive recovery in large-scale cloud AI systems, this paper proposes an intelligent self-healing framework that synergistically integrates large language models (LLMs) and deep reinforcement learning (DRL). The method employs LLM-driven environment modeling and action-space abstraction to jointly optimize fault semantic understanding and recovery policy generation. It further introduces a memory-guided meta-controller to enable continual adaptation to unseen fault patterns while mitigating catastrophic forgetting. Generalization is enhanced via multi-source log semantic parsing, prompt fine-tuning, and DRL experience replay. Evaluated on a cloud fault-injection platform, the framework reduces mean time to recovery by 37% under previously unseen fault scenarios—significantly outperforming state-of-the-art DRL and rule-based baselines. Key contributions include: (1) the first LLM-enabled environment modeling and action abstraction for joint fault semantics and recovery optimization; and (2) a memory-augmented meta-controller supporting robust continual learning in dynamic cloud environments.

Technology Category

Machine Learning: Life-Long and Continual LearningSearch and Optimization: Learning to SearchNatural Language Processing: (Large) Language Models

Application Category

Search and Retrieval-Augmented AI: Large language models for searchSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
As the scale and complexity of cloud-based AI systems continue to increase, the detection and adaptive recovery of system faults have become the core challenges to ensure service reliability and continuity. In this paper, we propose an Intelligent Fault Self-Healing Mechanism (IFSHM) that integrates Large Language Model (LLM) and Deep Reinforcement Learning (DRL), aiming to realize a fault recovery framework with semantic understanding and policy optimization capabilities in cloud AI systems. On the basis of the traditional DRL-based control model, the proposed method constructs a two-stage hybrid architecture: (1) an LLM-driven fault semantic interpretation module, which can dynamically extract deep contextual semantics from multi-source logs and system indicators to accurately identify potential fault modes; (2) DRL recovery strategy optimizer, based on reinforcement learning, learns the dynamic matching of fault types and response behaviors in the cloud environment. The innovation of this method lies in the introduction of LLM for environment modeling and action space abstraction, which greatly improves the exploration efficiency and generalization ability of reinforcement learning. At the same time, a memory-guided meta-controller is introduced, combined with reinforcement learning playback and LLM prompt fine-tuning strategy, to achieve continuous adaptation to new failure modes and avoid catastrophic forgetting. Experimental results on the cloud fault injection platform show that compared with the existing DRL and rule methods, the IFSHM framework shortens the system recovery time by 37% with unknown fault scenarios.
Problem

Research questions and friction points this paper is trying to address.

Detect and recover faults in cloud AI systems
Integrate LLM and DRL for fault semantic understanding
Improve exploration efficiency and generalization in reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates LLM and DRL for fault recovery
Uses LLM for semantic fault interpretation
Employs DRL for dynamic recovery optimization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Ze Yang
University of Illinois Urbana-Champaign, Champaign, USA
Yihong Jin
Yihong Jin
University of Illinois at Urbana-Champaign
Machine LearningPrivacy
J
Juntian Liu
Computer Science Department, University of Illinois Urbana-Champaign, Champaign, USA
X
Xinhe Xu
Computer Science Department, University of Illinois Urbana-Champaign, Champaign, USA