From Attack Simulation to SIEM Rule: Deterministic Detection-as-Code Synthesis with Probe-Level Traceability

πŸ“… 2026-06-03
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF

career value

189K/year
πŸ€– AI Summary
This study addresses the inefficiency and error-proneness of manually translating threats identified by Breach and Attack Simulation (BAS) into SIEM detection rules. To overcome this limitation, the authors propose a deterministic synthesis method that automatically maps BAS outputs to Sigma rules using a fixed corpus of probes, while preserving a complete, typed provenance chain from alerts back to their original probes. The approach leverages only 23 templates categorized according to the OWASP LLM/Web Top 10 and annotated with MITRE ATT&CK identifiers to generate byte-level stable, verifiable Sigma rules compatible with both Splunk and Elasticsearch. Evaluated on LLM and web probe corpora, the method successfully produced valid rules for all probes; on subsets of AdvBench and HarmBench, these LLM-focused rules triggered detections for 30% and 14% of attacks, respectively, with a false positive rate of 7.7%.
πŸ“ Abstract
Security teams routinely simulate attacks against their own systems to check whether their monitoring would catch a real intruder. These Breach-and-Attack-Simulation (BAS) tools surface findings, but the security information and event management (SIEM) systems that watch production need detection rules -- and today a human bridges that gap by hand, reading each finding and writing the corresponding Sigma rule (a vendor-neutral detection format). We show this translation can be partially automated when probes are drawn from a locked corpus, so each finding carries a stable identifier back to the originating probe. We describe a deterministic synthesis function that maps each finding to a starter Sigma rule through a small template library (N=23, indexed by categories from the OWASP LLM and Web Top 10), with a back-reference to the originating finding and its MITRE ATT&CK technique. On two locked corpora (17-probe LLM, 23-probe Web), every bypassed-probe finding yields a starter rule, and all 17/17 emitted rules parse and convert to Splunk and Elasticsearch backends. Replayed through a live OpenSearch SIEM, the LLM rules fire on 30% of a held-out AdvBench subset and 14% of HarmBench at 7.7% false positives on a benign baseline; the Web side is validated structurally, not against a held-out attack set. The contribution is a verifiable, byte-stable path from BAS finding to operator-deployable starter rule, re-derivable from the published corpus and template library alone -- trading the breadth of LLM-generative methods for exact reproducibility and a typed traceback from any fired alert to the originating probe.
Problem

Research questions and friction points this paper is trying to address.

Breach-and-Attack Simulation
SIEM
Detection Rule
Sigma Rule
Attack Traceability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Detection-as-Code
Breach-and-Attack Simulation
Sigma Rule Synthesis
Probe-Level Traceability
Deterministic Rule Generation
πŸ”Ž Similar Papers
No similar papers found.