From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high false-negative rates and insufficient investigation caused by reasoning deficiencies in LLM agents performing SOC alert triage. To overcome these limitations, this work proposes AIDA, a multi-agent framework that introduces an adversarial challenge mechanism, an append-only investigation ledger, and context-separated review. By employing dialectical analysis to reinforce evidence retrieval and decision verification, AIDA effectively mitigates single-agent cognitive biases. Experimental results demonstrate that AIDA achieves an F1 score of 0.958 and significantly reduces the false-negative rate from 40.4% to 3.1%. These findings indicate that the proposed framework substantially outperforms existing baseline methods while markedly decreasing the need for manual escalation in security operations.
📝 Abstract
Security operations centers (SOCs) must triage large volumes of alerts, most of which are benign, while missed attacks can remain uninvestigated. Tool-using large language model (LLM) agents can retrieve evidence during triage, but it remains unclear how reasoning strategies determine what to gather and when an investigation is sufficient to close an alert. We study five representative approaches spanning single-pass tool use, iterative retrieval, sampled investigations, self-review, and explicit verification. To support this study, we build ALERT-BENCH, an interactive benchmark that replays enterprise telemetry through a live SIEM and requires each system to retrieve evidence. Across 1,247 alerts from a multi-stage attack scenario, every approach missed at least 40.4% of attack-related alerts. Trace analysis shows that attack alerts are more likely to be dismissed when searches return no records, same-context review has negative net correction, and dismissal receives no consistently stronger investigation than escalation. Based on these findings, we further design AIDA (Adversarial Investigation and Dialectical Analysis), a multi-agent framework that requires an explicit proposed decision before independent challenge and stronger evidentiary requirements before dismissal. AIDA preserves investigation history in an append-only Investigation Ledger and keeps the challenge in a separate reasoning context. A separate Judge adjudicates the proposed decision and challenge against evidence, resolving the alert or requesting another round when evidence is missing. On the same alerts, AIDA achieves an F1 score of 0.958, compared with 0.371-0.744 for the studied approaches, and reduces the false-negative rate from 40.4% to 3.1% while escalating 18.4% of alerts to analysts. These results show that structuring evidence retrieval and decision review can substantially improve agentic SOC triage.
Problem

Research questions and friction points this paper is trying to address.

Alert Triage
SOC Agents
Large Language Models
False Negatives
Evidence Retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-agent framework
Alert triage
Adversarial investigation
Interactive benchmark
Large language model agents
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.