build automated remediation

Design, build, and evaluate systems and workflows that automatically generate, prioritize, execute, and track remediation actions for incidents, failures, or identified risks, including playbooks, pipelines, orchestration, and on-call/autonomous execution. Ensure safe automated fixes through prioritization, rollback and safety checks, remediation strategy and planning, and integration with incident-management and tracking mechanisms to verify resolution.

buildautomatedremediation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
2.8
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$195K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of costly erroneous repairs in existing automated remediation systems, which often lack the ability to assess intervention necessity and thus rely on manual approval for safety. The authors formulate safe repair as an intervention decision problem under risk constraints and introduce a three-dimensional risk decomposition framework encompassing impact scope, reversibility, and epistemic uncertainty. They further design a context-adaptive human-in-the-loop gating strategy that enables interpretable, workload-aware safety interventions. Built upon constrained Markov decision processes (CMDPs), offline policy learning, Chaos Mesh fault injection, and the RCAEval classification framework, the proposed approach reduces erroneous repair rates by 39% and improves repair success rates by 2.5 percentage points on the Train Ticket benchmark, while decreasing on-call escalation burden by 17% compared to fixed-threshold baselines.

automated decision-makingfalse remediation ratemicroservice systems

Towards an Engineering Workflow Management System for Asset Administration Shells using BPMN

Jul 10, 2025
SG
Sten Grüner
🏛️ Process Control Platform | ABB AG | ABB AG Corporate Research Center

To address the insufficient security and scalability of engineering workflow automation and cross-organizational collaboration in Industry 4.0, this paper proposes an engineering workflow management approach integrating Asset Administration Shells (AAS) with BPMN. We innovatively design a distributed, write-on-copy AAS infrastructure to ensure data consistency and access security, and develop a lightweight workflow engine prototype supporting native AAS operations, enabling automatic mapping and execution of BPMN processes onto AAS interactions. This method unifies digital twin representation, asset modeling, and business process logic, thereby significantly enhancing standardization of engineering data exchange, end-to-end process traceability, and multi-stakeholder collaboration efficiency. Experimental evaluation demonstrates the system’s feasibility for secure inter-organizational coordination and its horizontal scalability across heterogeneous industrial environments.

Automate AAS operations and engineering workflows efficientlyEnhance security and scalability of Asset Administration ShellsIntegrate Industry 4.0 technologies into engineering workflows

Agentic Troubleshooting Guide Automation for Incident Management

Oct 11, 2025
JM
JIAYI MAO
🏛️ Tsinghua University | Microsoft | Microsoft Research

Manual execution of Troubleshooting Guides (TSGs) in large-scale IT systems is inefficient and error-prone, while existing LLM-based approaches struggle with poor TSG quality, complex control flow, data-intensive queries, and parallel execution requirements. Method: We propose an end-to-end automation framework comprising: (i) TSG Mentor to enhance guide quality; (ii) an offline phase leveraging LLMs to construct a structured execution DAG and generate domain-specific Query Preparation Plugins (QPPs); and (iii) an online phase employing a DAG-guided, memory-augmented scheduler and executor that ensures correctness and enables task-level parallelism. Results: Evaluated on real-world TSGs and incidents, our framework achieves a 94% success rate with GPT-4.1—significantly outperforming baselines—and reduces execution time for parallelizable TSGs by 32.9%–70.4%, while also improving token efficiency and latency.

Addressing LLM limitations in handling complex control flow and data queriesAutomating troubleshooting guides to reduce manual execution errors and delaysImproving parallel execution efficiency for IT incident management workflows

This work proposes a novel end-to-end paradigm for automated microservice repair that overcomes key limitations of existing large language model (LLM)-based approaches, which often rely on handcrafted prompts, lack runtime contextual knowledge, and suffer from the accuracy and efficiency constraints of general-purpose models. The proposed method directly generates executable Ansible playbooks from diagnostic reports and introduces MicroRemed, a comprehensive benchmark enabling automated deployment, fault injection, and repair validation. By leveraging empirically simulated data to perform reinforcement fine-tuning, the approach trains a specialized repair model that eliminates dependence on expert-crafted prompts and generic LLMs. Experimental results demonstrate that this method significantly outperforms nine representative LLMs on both public and industrial microservice platforms, achieving substantial improvements in both repair accuracy and execution efficiency.

auto-remediationdiagnosis reportsexecutable playbooks

Latest Papers

What's happening recently
View more

This work addresses the limitations of traditional expert-manual-based cybersecurity response methods, which struggle to adapt to dynamic attack scenarios and evolving recovery objectives, as well as the instability of existing large-model approaches in long-horizon tasks. The authors propose an end-to-end agent planning framework that innovatively models event states using a graph structure (Graph-as-State), incorporates a phase-aware agent routing mechanism, and establishes a verifiable experience reuse loop to guide action selection and state updates. The system integrates multi-agent large language models with experience retrieval augmentation and execution feedback verification, enabling dynamic, stable, and evolvable response planning within a Docker-based network range simulation environment. Experimental results demonstrate that the proposed method achieves a normalized defense score of 0.94 across 100 simulated scenarios, representing a 9.5% improvement over the strongest baseline.

adaptive responseagentic planningcyberattack recovery

This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).

Agentic SystemsMonitoringStructural Defects

This work addresses the limitations of existing large language models, which are typically confined to isolated tasks and struggle to integrate into industrial-scale, multi-stage security workflows. To bridge this gap, the authors propose the first role-based multi-agent framework tailored to the entire vulnerability lifecycle, incorporating specialized agents—Planner, Analyzer, Fixer, and Verifier—augmented with CodeQL static analysis for enhanced precision. By introducing a role-oriented multi-agent architecture into end-to-end vulnerability management, this approach effectively aligns the capabilities of large models with real-world security engineering demands. Evaluated on 25 real-world C/C++ vulnerabilities, the system achieves a detection accuracy of 44%—comparable to GPT-5.5—and a repair accuracy of 19%, offering a practical and collaborative paradigm for intelligent security operations.

LLM-based securityrole-based agentic architecturesecure software engineering

This study addresses the challenge of automating workflows in complex industries—such as logistics, healthcare, and construction—where processes are fragmented across heterogeneous tools and involve multi-party collaboration. The work proposes orchestration as a core abstraction to enable effective automation by dynamically coordinating multi-step tasks, enforcing domain-specific constraints, managing human approvals, and integrating legacy systems. It introduces the novel concept of “orchestration bottlenecks” and develops a theoretical framework that unifies multi-agent systems, workflow modeling, constraint reasoning, and human–AI collaboration, while exposing critical gaps in current multi-agent approaches at the orchestration level. Based on distinct sources of operational friction across domains, the paper advocates for targeted architectural safeguards—such as constraint enforcement or explainability—and phased implementation strategies to provide actionable pathways for automation in complex operational environments.

legacy systemsoperationally complex industriesorchestration

This study addresses the limitations of existing autonomous business process execution approaches, which predominantly focus on control-flow constraints and struggle to support compliance-aware decision-making under multifaceted requirements involving data-aware and temporal conditions. To overcome this gap, the work introduces a unified multi-perspective framework that formally integrates data and time constraints through a numeric planning-based modeling approach. This enables efficient what-if analysis and optimal continuation recommendations for partially executed processes. Experimental results demonstrate that the proposed method not only ensures regulatory compliance but also exhibits strong scalability, substantially enhancing the effectiveness and practicality of autonomous decision-making in AI-augmented business process management systems.

Business Process ManagementFramed AutonomyMulti-Perspective Constraints

Hot Scholars

YL

Yuekang Li

Lecturer (Assistant Professor), University of New South Wales
Software EngineeringSoftware SecurityAI Red Teaming
GD

Gelei Deng

Nanyang Technological University
CybersecuritySystem securityRobotics SecurityAI Security
YL

Yi Liu

AI Research @ Quantstamp | PhD @ NTU | BEng @ SUSTech
AI AgentSoftware EngineeringLLM Security
LM

Lei Ma

Associate Professor, The University of Tokyo, Japan & University of Alberta, Canada
Trustworthy AIMLOpsSE4AIAI Safety
YZ

Ying Zhang

Wake Forest University
Static programming analysisVulnerable API detectionSecurity test generation