RuleGenie: SIEM Detection Rule Set Optimization

📅 2025-05-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
SIEM rule redundancy leads to high false-positive rates, analyst fatigue, and delayed incident response; existing manual optimization approaches suffer from low efficiency and poor scalability. This paper introduces the first LLM-driven automated rule-set optimization framework, integrating Transformer-based multi-head attention embeddings with semantic similarity matching to enable unified modeling and compression of rules across heterogeneous platforms—including Splunk, Sigma, and AQL. Leveraging large language models’ capabilities in information extraction, logical reasoning, and natural language understanding, the framework performs semantic-level deduplication and generates actionable optimization recommendations. Evaluated on real-world enterprise-scale rule sets, it significantly reduces false positives, enhances rule-set compactness and execution efficiency, and demonstrates strong generalizability and platform independence.

Technology Category

Search and Optimization: Algorithm ConfigurationData Mining & Knowledge Management: Rule Mining & Pattern MiningMachine Learning: Large Multimodal Models (LMMs)

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Large language models for searchUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
SIEM systems serve as a critical hub, employing rule-based logic to detect and respond to threats. Redundant or overlapping rules in SIEM systems lead to excessive false alerts, degrading analyst performance due to alert fatigue, and increase computational overhead and response latency for actual threats. As a result, optimizing SIEM rule sets is essential for efficient operations. Despite the importance of such optimization, research in this area is limited, with current practices relying on manual optimization methods that are both time-consuming and error-prone due to the scale and complexity of enterprise-level rule sets. To address this gap, we present RuleGenie, a novel large language model (LLM) aided recommender system designed to optimize SIEM rule sets. Our approach leverages transformer models' multi-head attention capabilities to generate SIEM rule embeddings, which are then analyzed using a similarity matching algorithm to identify the top-k most similar rules. The LLM then processes the rules identified, utilizing its information extraction, language understanding, and reasoning capabilities to analyze rule similarity, evaluate threat coverage and performance metrics, and deliver optimized recommendations for refining the rule set. By automating the rule optimization process, RuleGenie allows security teams to focus on more strategic tasks while enhancing the efficiency of SIEM systems and strengthening organizations' security posture. We evaluated RuleGenie on a comprehensive set of real-world SIEM rule formats, including Splunk, Sigma, and AQL (Ariel query language), demonstrating its platform-agnostic capabilities and adaptability across diverse security infrastructures. Our experimental results show that RuleGenie can effectively identify redundant rules, which in turn decreases false positive rates and enhances overall rule efficiency.
Problem

Research questions and friction points this paper is trying to address.

Optimizing SIEM rule sets to reduce false alerts and computational overhead
Automating rule optimization using LLM to replace manual error-prone methods
Enhancing SIEM efficiency by identifying redundant rules across diverse platforms
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-aided recommender for SIEM rule optimization
Transformer models generate SIEM rule embeddings
Similarity matching and LLM analyze rule sets
🔎 Similar Papers
No similar papers found.