🤖 AI Summary
This study addresses the challenge of LLM agents adapting to dynamic organizational standards during alert triage in Security Operations Centers. To this end, it proposes a feedback-driven continuous evolution framework grounded in structured skill representations. The core innovation lies in enforcing a hard constraint of 1.0 recall to maximize the automated closure of false positives, while identifying judgment blind spots by integrating alert distributions with model error boundaries. Experimental evaluations across four real-world industrial scenarios demonstrate that the proposed method achieves full recall on all evolution sets, with three scenarios maintaining perfect recall throughout future testing windows. These results indicate that the framework significantly outperforms baseline approaches, offering a robust and adaptive solution for automated security alert triage under evolving operational criteria.
📝 Abstract
Security Operations Centers (SOCs) process large volumes of alerts daily. Alert triage prioritizes high-risk threats while reducing manual review of benign alerts. LLM agents can reason over logs and threat intelligence, but struggle to keep aligned with organization-specific, rapidly evolving SOC operational standards. We introduce REFINE, an LLM-agent framework for enterprise alert triage. REFINE encodes analyst expertise as structured skills and continuously adapts using analyst disposition feedback. It enforces recall = 1.0 as a hard constraint during evolution to maximize auto-closure of false positives, and identifies judgment blind spots by combining alert distributions with model error boundaries. Evaluated on four real industrial SOC scenarios across four MITRE ATT&CK phases with temporal split: REFINE achieves recall=1.0 on all evolution sets. On future test windows, it retains recall=1.0 in three scenarios; the degraded case reaches 0.807 recall, still outperforming self-evolution baselines (0.49-0.58).