respond to incidents

Designs, implements, and operates the processes, playbooks, tools, and automations used to detect, triage, coordinate, troubleshoot, investigate, and remediate incidents; defines escalation paths, on-call rotations, and incident response workflows for consistent handling. Performs incident debugging, root-cause and post‑incident analysis, documents findings and lessons learned, and updates response planning and automation to prevent recurrence and improve response effectiveness.

respondtoincidents

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$194K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of traditional expert-manual-based cybersecurity response methods, which struggle to adapt to dynamic attack scenarios and evolving recovery objectives, as well as the instability of existing large-model approaches in long-horizon tasks. The authors propose an end-to-end agent planning framework that innovatively models event states using a graph structure (Graph-as-State), incorporates a phase-aware agent routing mechanism, and establishes a verifiable experience reuse loop to guide action selection and state updates. The system integrates multi-agent large language models with experience retrieval augmentation and execution feedback verification, enabling dynamic, stable, and evolvable response planning within a Docker-based network range simulation environment. Experimental results demonstrate that the proposed method achieves a normalized defense score of 0.94 across 100 simulated scenarios, representing a 9.5% improvement over the strongest baseline.

adaptive responseagentic planningcyberattack recovery

Agentic Troubleshooting Guide Automation for Incident Management

Oct 11, 2025
JM
JIAYI MAO
🏛️ Tsinghua University | Microsoft | Microsoft Research

Manual execution of Troubleshooting Guides (TSGs) in large-scale IT systems is inefficient and error-prone, while existing LLM-based approaches struggle with poor TSG quality, complex control flow, data-intensive queries, and parallel execution requirements. Method: We propose an end-to-end automation framework comprising: (i) TSG Mentor to enhance guide quality; (ii) an offline phase leveraging LLMs to construct a structured execution DAG and generate domain-specific Query Preparation Plugins (QPPs); and (iii) an online phase employing a DAG-guided, memory-augmented scheduler and executor that ensures correctness and enables task-level parallelism. Results: Evaluated on real-world TSGs and incidents, our framework achieves a 94% success rate with GPT-4.1—significantly outperforming baselines—and reduces execution time for parallelizable TSGs by 32.9%–70.4%, while also improving token efficiency and latency.

Addressing LLM limitations in handling complex control flow and data queriesAutomating troubleshooting guides to reduce manual execution errors and delaysImproving parallel execution efficiency for IT incident management workflows

DrP: Meta's Efficient Investigations Platform at Scale

Dec 03, 2025
SS
Shubham Somani
🏛️ Meta

In large-scale systems, on-call engineers rely on manual procedures or ad-hoc scripts for incident investigation, resulting in high mean time to resolution (MTTR), elevated operational overhead, and diminished productivity. This paper introduces DrP—the first end-to-end automated investigation framework designed for heterogeneous domains including services, AI/ML, and mobile systems. DrP’s key contributions are: (1) a declarative SDK enabling low-code development of reusable, domain-agnostic analysis logic; (2) a distributed execution engine with a plugin-based architecture supporting high-concurrency diagnostics and deep integration with alerting, event management, and remediation systems; and (3) a unified abstraction layer that transparently insulates users from infrastructure heterogeneity. Deployed at scale within Meta, DrP executes ~50,000 analyses daily across 300+ engineering teams, reducing average MTTR by 20% overall and up to 80% in specific scenarios—significantly enhancing SRE responsiveness and system observability.

Automates manual investigation processes to reduce incident resolution timeProvides an end-to-end framework for scalable, automated incident analysis and mitigationReduces on-call toil and improves productivity in large-scale systems

Incident Analysis for AI Agents

Aug 19, 2025
CE
Carson Ezell
🏛️ Harvard University | Centre for the Governance of AI

Current AI agent incident reporting mechanisms suffer from critical limitations: they rely solely on publicly available data, omitting sensitive yet essential internal execution traces—such as reasoning chains and tool invocation logs—thereby hindering identification of root causes. This work pioneers the application of systems safety principles to AI agent incident analysis, proposing a novel “Systemic–Contextual–Cognitive” tri-dimensional causal framework. We design an integrated attribution methodology combining activity log analysis, system documentation review, and tool behavior tracing. Furthermore, we specify mandatory fields for incident reports and define a minimal sensitive dataset that developers and deployers must retain—including reasoning trajectories, API call sequences, and environmental context. Our contributions establish both a theoretical foundation and practical guidelines for building explainable, reproducible, and intervenable AI agent incident response mechanisms. (149 words)

Addressing insufficient incident reporting processes for AI agent failuresAnalyzing AI agent incidents to understand causes and prevent harmProposing a framework to identify system, contextual, and cognitive factors

Latest Papers

What's happening recently
View more

This work addresses inefficiencies in Network Operations Centers (NOCs)—including fragmented information retrieval, verbose ticketing, and loss of contextual continuity during handoffs—stemming from data silos. To mitigate these challenges, the authors propose ORBIT, an intelligent agent system integrated into the ServiceNow platform. ORBIT employs a modular, layered architecture that encapsulates task logic into versioned, testable “skills,” ensuring reliable and scalable operation within constrained behavioral boundaries. Its core components comprise a centralized reasoning engine, an MCP protocol interface to ESnet, a semantic search layer, an operational chat interface, and a LiteLLM model gateway. Evaluated on six initial tasks and rapidly adapted to two new ones, ORBIT significantly reduces operational steps, eliminates known error patterns, and sees broad adoption of its reusable components, thereby lowering cognitive load and accelerating incident response.

cognitive loadcontext lossincident resolution

为解决ERP系统中数据集成和流程监控的碎片化问题,本文提出一种企业流程控制塔,通过集成状态观测、语义翻译、机器学习诊断等方法提升IT团队的工作效率。

Electronic Data Interchange (EDI)Enterprise Resource Planning (ERP)Intermediate Document (IDoc)

This study addresses the high false-negative rates and insufficient investigation caused by reasoning deficiencies in LLM agents performing SOC alert triage. To overcome these limitations, this work proposes AIDA, a multi-agent framework that introduces an adversarial challenge mechanism, an append-only investigation ledger, and context-separated review. By employing dialectical analysis to reinforce evidence retrieval and decision verification, AIDA effectively mitigates single-agent cognitive biases. Experimental results demonstrate that AIDA achieves an F1 score of 0.958 and significantly reduces the false-negative rate from 40.4% to 3.1%. These findings indicate that the proposed framework substantially outperforms existing baseline methods while markedly decreasing the need for manual escalation in security operations.

Alert TriageEvidence RetrievalFalse Negatives

This work addresses the inefficiency of traditional manual security incident response and the limitations of existing automated approaches, which are either difficult to deploy or suffer from unreliable planning due to hallucinations when relying solely on large language models (LLMs). To overcome these challenges, the paper proposes a novel multi-scale intelligent response architecture that integrates decision-theoretic planning with a lightweight LLM. The framework uniquely combines digital twins, a tactical-level rollout planner, and an operational-level LLM agent, establishing a dual-scale (tactical–operational) coordination mechanism to enable reliable and executable automated responses in simulated environments. Experimental results across three attack scenarios demonstrate that the proposed approach reduces average recovery time by 15.1% and improves success rate by 33.6% compared to state-of-the-art LLM-based baselines.

automated planningdecision-theoretic planningdigital twin

Hot Scholars

SC

Stephen Casper

PhD student, MIT
AI safetyAI responsibilityred-teamingrobustness
GD

Gelei Deng

Nanyang Technological University
CybersecuritySystem securityRobotics SecurityAI Security
NE

Nelly Elsayed

University of Cincinnati
Applied AIHealthcare InformaticsCybersecurityIntelligent Information Systems
YZ

Yajin Zhou

Zhejiang University
Blockchain System Security
FK

Foutse Khomh

NSERC Arthur B. McDonald Fellow, CRC Tier 1, Canada CIFAR AI Chair, FRQ-IVADO Chair, Full Professor
Software engineeringMachine learning systems engineeringMining software repositoriesReverse