Sigma-Hunter: A Domain-Specific Language Model for Threat Hunting and Detection Engineering

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of general-purpose large language models in generating Sigma rules, which frequently exhibit syntax errors, missing fields, and overly broad logic. Leveraging a validated open-source Sigma rule dataset, this work employs Low-Rank Adaptation (LoRA) to instruction-tune Mistral and Phi-4, thereby constructing domain-specific models tailored for threat hunting. The findings demonstrate that 7B-parameter models, when properly adapted to the target domain, achieve performance comparable to larger counterparts. Furthermore, a notable divergence between syntactic validity and semantic quality is observed. The proposed model attains a composite score of 8.17, outperforming the strongest baseline while supporting local offline deployment. Ultimately, this approach significantly enhances the semantic accuracy of automated Sigma rule generation for cybersecurity applications.
📝 Abstract
Detection engineers must translate threat reports, forensic observations, and hunt hypotheses into precise, testable rules. General-purpose large language models (LLMs) can draft such rules, but often produce invalid YAML, incorrect log sources, unsupported fields, or overly broad detection logic. This paper presents \emph{Sigma-Hunter}, a domain-adapted LLM for analyst-assistive Sigma rule generation and threat hunting. We build an instruction-tuning dataset from 3,635 validated open-source Sigma rules, expanded into 7,663 question-answer and analyst-reasoning examples. Each source rule is assigned to a single train, validation, or test partition before this expansion, so no rule leaks across splits. We fine-tune a 7B Mistral model and a Phi-4 model with LoRA and score held-out rule generations on syntax, approximate field consistency, and a semantic judgment of detection logic, completeness, selectivity, and log-source alignment. Sigma-Hunter-Mistral scores 8.17 overall, against 7.88 for the strongest general-purpose baseline and 4.61 for untuned Mistral. Two findings stand out: domain adaptation enables a compact 7B model to perform competitively with larger general-purpose models on this structured task, and syntactic validity is a weak proxy for semantic rule quality, as several baselines emit well-formed YAML carrying weak detection logic. The adapted models run locally, which suits detection engineering in disconnected environments where analysts cannot reach hosted model services.
Problem

Research questions and friction points this paper is trying to address.

Threat Hunting
Detection Engineering
Sigma Rules
Domain-Specific LLM
Rule Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Domain-Specific Language Model
Sigma Rule Generation
Instruction Tuning
LoRA Fine-tuning
Threat Hunting