🤖 AI Summary
本文针对DevOps部署中的自动故障恢复难题,提出了一种结合基于规则和机器学习的混合框架,通过实时监控、识别、分类并恢复故障,有效提高了系统的恢复效率与运行稳定性。
📝 Abstract
Automatic failure recovery is another difficult area of DevOps deployments due to the complexity of the system, high workload and the limitations of the traditional rules-based or manual approach. This study proposes a hybrid model integrating deterministic, rule-based recovery with the assistance of machine learning by fault prediction to automatically monitor failures, identify, classify and recover failures in real time. To test the architecture, a distributed log dataset of 100,000 records was used for models such as Decision Tree, Random Forest, Logistic Regression, LightGBM, Autoencoder with BiLSTM. The Decision Tree had a strong performance in defect identification with an F1 score of 80.8%, recall of 80.0%, precision of 81.6%, and accuracy of 89.4%. An impressive 83.3% success rate was achieved by the automated recovery measures, resulting in a 94.6% decrease in mean time to recovery (MTTR), a 95.4% reduction in downtime, and annual savings of 7,462. The combination of explainable rule-based logic and adaptive AI prediction in the framework means it is more resilient, operationally efficient, and scalable, and an effective, autonomous approach to DevOps environments evaluated under controlled experimental conditions.