Score
Designs and implements integrations between SIEM platforms and telemetry/log sources, including connector configuration, log ingestion pipelines, parsing, and normalization. Builds and tunes correlation rules and alerting logic, performs SIEM log analysis to reduce false positives and optimize detections, and configures or integrates specific SIEM products (e.g., Microsoft Sentinel).
This study addresses the inefficiency and error-proneness of manually translating threats identified by Breach and Attack Simulation (BAS) into SIEM detection rules. To overcome this limitation, the authors propose a deterministic synthesis method that automatically maps BAS outputs to Sigma rules using a fixed corpus of probes, while preserving a complete, typed provenance chain from alerts back to their original probes. The approach leverages only 23 templates categorized according to the OWASP LLM/Web Top 10 and annotated with MITRE ATT&CK identifiers to generate byte-level stable, verifiable Sigma rules compatible with both Splunk and Elasticsearch. Evaluated on LLM and web probe corpora, the method successfully produced valid rules for all probes; on subsets of AdvBench and HarmBench, these LLM-focused rules triggered detections for 30% and 14% of attacks, respectively, with a false positive rate of 7.7%.
SIEM rule redundancy leads to high false-positive rates, analyst fatigue, and delayed incident response; existing manual optimization approaches suffer from low efficiency and poor scalability. This paper introduces the first LLM-driven automated rule-set optimization framework, integrating Transformer-based multi-head attention embeddings with semantic similarity matching to enable unified modeling and compression of rules across heterogeneous platforms—including Splunk, Sigma, and AQL. Leveraging large language models’ capabilities in information extraction, logical reasoning, and natural language understanding, the framework performs semantic-level deduplication and generates actionable optimization recommendations. Evaluated on real-world enterprise-scale rule sets, it significantly reduces false positives, enhances rule-set compactness and execution efficiency, and demonstrates strong generalizability and platform independence.
This work addresses the scarcity of production-grade Security Operations Center (SOC) logs for research due to stringent privacy constraints, which has led prior studies to rely on synthetic or outdated data. To bridge this gap, the authors propose a methodology that, for the first time, transforms real-world financial-sector SIEM logs into reusable research artifacts while adhering to strict privacy boundaries. The approach preserves investigation-relevant structures through structured anonymization, mapping to the MITRE ATT&CK framework, deterministic validation, and large language model (LLM)-based behavioral compliance checks. The resulting artifact comprises 37 HIKARI challenges suitable for effective model training and demonstrates its utility by accurately identifying LLM policy violations across 200 SOCpilot incidents, thereby validating its balanced trade-off between privacy preservation and analytical fidelity.
Manual mapping of SIEM rules to MITRE ATT&CK techniques (TTPs) is inefficient and error-prone, while existing machine learning approaches struggle to handle the structured nature of SIEM rules. Method: We propose the first multi-stage prompt-chaining LLM framework specifically designed for SIEM rule-to-TTP mapping—requiring no pretraining or fine-tuning. It leverages structured prompt engineering and injection of external cybersecurity knowledge to enhance domain-specific TTP identification capabilities of large language models (e.g., GPT-4-Turbo, Qwen, Granite, Mistral). Contribution/Results: Evaluated on the Splunk Security Content dataset, GPT-4-Turbo achieves the highest accuracy. Ablation studies confirm that external knowledge substantially compensates for LLMs’ deficiencies in implicit cybersecurity knowledge. This work establishes a new paradigm for automated, interpretable, and scalable annotation of threat-detection rules.
Existing log analysis models are task-specific, rely heavily on domain-specific annotated data, exhibit poor generalization, and struggle with complex or unseen instructions. Method: We propose LogLM, an instruction-driven large language model for log analysis, which unifies diverse log tasks—including anomaly detection, parsing, and summarization—into a standardized instruction-response format. LogLM is adapted to the log domain via multi-task instruction tuning and log-specific instruction engineering. It accepts natural-language instructions and supports zero-shot cross-task transfer. Contribution/Results: Experiments demonstrate that LogLM outperforms all state-of-the-art methods across five core log analysis tasks. It exhibits strong generalization to complex instructions and previously unseen tasks. As a single unified model, LogLM replaces multiple specialized models, significantly improving deployment efficiency and task-agnostic capability.
This work addresses the absence of an organization-level agent runtime architecture in financial cybersecurity workflows that supports both model-agnostic operation and on-premises deployment, thereby hindering consistent enforcement of security policies across retrieval, tool invocation, and auditing stages. To bridge this gap, the paper proposes a novel architecture featuring a typed security context propagated throughout the entire workflow, integrated with SIEM/XDR systems as contextual data sources. The design incorporates a managed tool adaptation layer, structured evidence referencing, and a hierarchical human-agent collaboration mechanism. Key innovations include a shared runtime core, logically specialized sub-agents, append-only auditing, and optional extensions such as graph-based retrieval and MCP protocol support. The study defines testable architectural slices and establishes a falsifiable evaluation framework encompassing policy enforcement, evidence traceability, output quality, and observability.
This work addresses the challenges of root cause diagnosis in large-scale microservice systems, where existing approaches are hindered by massive log volumes, limited LLM context windows, and insufficient semantic reasoning and interpretability. The authors propose a neuro-symbolic hybrid method that emulates Site Reliability Engineers’ manual troubleshooting process through a six-stage pipeline for log sampling, template clustering, and anomaly ranking, producing a concise evidence package for LLM-based root cause inference. This approach compresses raw logs by 1,000–7,000× while preserving critical failure signals and provides auditable log templates and statistical evidence, substantially enhancing interpretability and practicality. Evaluated on 11 real-world incidents, the method achieves an MRR of 0.790 and ranks the correct root cause within the top three candidates in over 90% of cases within one minute, earning strong endorsement from operations teams.
This work addresses the pervasive issue of redundant and isolated messages in system logs, which hinder downstream tasks such as model reasoning and anomaly detection. To tackle this challenge, the authors propose LogPurifier—the first task-agnostic log cleansing framework—that systematically purifies logs by extracting log templates and modeling their dependencies to accurately identify and remove messages irrelevant to system functional behavior. By doing so, LogPurifier enables effective log sanitization applicable across diverse analytical scenarios. Experimental results demonstrate that LogPurifier substantially improves both accuracy and efficiency in various downstream tasks, thereby validating its effectiveness and generalizability.