Score
Designs, implements, and operates the processes, playbooks, tools, and automations used to detect, triage, coordinate, troubleshoot, investigate, and remediate incidents; defines escalation paths, on-call rotations, and incident response workflows for consistent handling. Performs incident debugging, root-cause and post‑incident analysis, documents findings and lessons learned, and updates response planning and automation to prevent recurrence and improve response effectiveness.
This work addresses the limitations of traditional expert-manual-based cybersecurity response methods, which struggle to adapt to dynamic attack scenarios and evolving recovery objectives, as well as the instability of existing large-model approaches in long-horizon tasks. The authors propose an end-to-end agent planning framework that innovatively models event states using a graph structure (Graph-as-State), incorporates a phase-aware agent routing mechanism, and establishes a verifiable experience reuse loop to guide action selection and state updates. The system integrates multi-agent large language models with experience retrieval augmentation and execution feedback verification, enabling dynamic, stable, and evolvable response planning within a Docker-based network range simulation environment. Experimental results demonstrate that the proposed method achieves a normalized defense score of 0.94 across 100 simulated scenarios, representing a 9.5% improvement over the strongest baseline.
Manual execution of Troubleshooting Guides (TSGs) in large-scale IT systems is inefficient and error-prone, while existing LLM-based approaches struggle with poor TSG quality, complex control flow, data-intensive queries, and parallel execution requirements. Method: We propose an end-to-end automation framework comprising: (i) TSG Mentor to enhance guide quality; (ii) an offline phase leveraging LLMs to construct a structured execution DAG and generate domain-specific Query Preparation Plugins (QPPs); and (iii) an online phase employing a DAG-guided, memory-augmented scheduler and executor that ensures correctness and enables task-level parallelism. Results: Evaluated on real-world TSGs and incidents, our framework achieves a 94% success rate with GPT-4.1—significantly outperforming baselines—and reduces execution time for parallelizable TSGs by 32.9%–70.4%, while also improving token efficiency and latency.
In large-scale systems, on-call engineers rely on manual procedures or ad-hoc scripts for incident investigation, resulting in high mean time to resolution (MTTR), elevated operational overhead, and diminished productivity. This paper introduces DrP—the first end-to-end automated investigation framework designed for heterogeneous domains including services, AI/ML, and mobile systems. DrP’s key contributions are: (1) a declarative SDK enabling low-code development of reusable, domain-agnostic analysis logic; (2) a distributed execution engine with a plugin-based architecture supporting high-concurrency diagnostics and deep integration with alerting, event management, and remediation systems; and (3) a unified abstraction layer that transparently insulates users from infrastructure heterogeneity. Deployed at scale within Meta, DrP executes ~50,000 analyses daily across 300+ engineering teams, reducing average MTTR by 20% overall and up to 80% in specific scenarios—significantly enhancing SRE responsiveness and system observability.
Current AI agent incident reporting mechanisms suffer from critical limitations: they rely solely on publicly available data, omitting sensitive yet essential internal execution traces—such as reasoning chains and tool invocation logs—thereby hindering identification of root causes. This work pioneers the application of systems safety principles to AI agent incident analysis, proposing a novel “Systemic–Contextual–Cognitive” tri-dimensional causal framework. We design an integrated attribution methodology combining activity log analysis, system documentation review, and tool behavior tracing. Furthermore, we specify mandatory fields for incident reports and define a minimal sensitive dataset that developers and deployers must retain—including reasoning trajectories, API call sequences, and environmental context. Our contributions establish both a theoretical foundation and practical guidelines for building explainable, reproducible, and intervenable AI agent incident response mechanisms. (149 words)
本文设计了一个基于规则引擎和本地大型语言模型的教育信息系统警报后事件协调与响应子系统,通过实验验证了其功能正确性和可控性。
This work addresses inefficiencies in Network Operations Centers (NOCs)—including fragmented information retrieval, verbose ticketing, and loss of contextual continuity during handoffs—stemming from data silos. To mitigate these challenges, the authors propose ORBIT, an intelligent agent system integrated into the ServiceNow platform. ORBIT employs a modular, layered architecture that encapsulates task logic into versioned, testable “skills,” ensuring reliable and scalable operation within constrained behavioral boundaries. Its core components comprise a centralized reasoning engine, an MCP protocol interface to ESnet, a semantic search layer, an operational chat interface, and a LiteLLM model gateway. Evaluated on six initial tasks and rapidly adapted to two new ones, ORBIT significantly reduces operational steps, eliminates known error patterns, and sees broad adoption of its reusable components, thereby lowering cognitive load and accelerating incident response.
为解决ERP系统中数据集成和流程监控的碎片化问题,本文提出一种企业流程控制塔,通过集成状态观测、语义翻译、机器学习诊断等方法提升IT团队的工作效率。
This study addresses the high false-negative rates and insufficient investigation caused by reasoning deficiencies in LLM agents performing SOC alert triage. To overcome these limitations, this work proposes AIDA, a multi-agent framework that introduces an adversarial challenge mechanism, an append-only investigation ledger, and context-separated review. By employing dialectical analysis to reinforce evidence retrieval and decision verification, AIDA effectively mitigates single-agent cognitive biases. Experimental results demonstrate that AIDA achieves an F1 score of 0.958 and significantly reduces the false-negative rate from 40.4% to 3.1%. These findings indicate that the proposed framework substantially outperforms existing baseline methods while markedly decreasing the need for manual escalation in security operations.
本文探讨了如何通过不同类型的演练来提高自动驾驶车辆事故管理的有效性,提出了一套针对自动驾驶车辆特性的演练框架。
This work addresses the inefficiency of traditional manual security incident response and the limitations of existing automated approaches, which are either difficult to deploy or suffer from unreliable planning due to hallucinations when relying solely on large language models (LLMs). To overcome these challenges, the paper proposes a novel multi-scale intelligent response architecture that integrates decision-theoretic planning with a lightweight LLM. The framework uniquely combines digital twins, a tactical-level rollout planner, and an operational-level LLM agent, establishing a dual-scale (tactical–operational) coordination mechanism to enable reliable and executable automated responses in simulated environments. Experimental results across three attack scenarios demonstrate that the proposed approach reduces average recovery time by 15.1% and improves success rate by 33.6% compared to state-of-the-art LLM-based baselines.