Score
Designs and implements pipelines, instrumentation, and parsers that extract, normalize, and structure telemetry (logs and SIEM streams) into canonical records—handling timestamp and event-field normalization, log parsing and normalization, enrichment, and collapsing correlated events into artifacts while preserving investigative relations and ordering. Builds and analyzes algorithms and procedures (e.g., canonical deletion scans and irredundant-core decomposition) that identify and remove redundant facts, compute a semantic core and remainder, and enable core-only rate reductions without breaking semantic closure.
Security logs are inherently unstructured, semantically ambiguous, and ill-suited for deep reasoning. Method: We propose an ontology-driven, large language model (LLM)-augmented knowledge graph construction framework. It integrates a cybersecurity ontology—aligned with NIS 2 and EU classification standards—with semantic log parsing and LLM-enhanced entity-relation extraction to automate log-to-knowledge-graph mapping; ontology constraints improve LLM output accuracy and interpretability. Contribution/Results: This work establishes the first end-to-end semantic–knowledge co-analytical paradigm for security logs. Experiments demonstrate an 18.7% improvement in log structuring F1-score and significantly enhanced threat contextual reasoning. The framework enables cross-system data interoperability and provides a verifiable semantic foundation for intelligent threat detection and regulatory compliance auditing.
This study addresses the challenge of performance evaluation and optimization of log parsers by systematically comparing syntactic versus semantic approaches and single-stage versus two-stage architectures. We propose SynLog+, a lightweight template identification enhancement module designed as the second stage of a two-stage parsing framework; it jointly leverages syntactic analysis and semantic modeling to significantly improve accuracy with negligible runtime overhead. Experiments across diverse benchmarks demonstrate that SynLog+ boosts average accuracy by 236% for syntactic parsers and by 20% for semantic parsers, confirming its superior accuracy–efficiency trade-off. Our core contributions are twofold: (1) the first generalizable, architecture-agnostic enhancement design for template identification within two-stage log parsing frameworks; and (2) a structured, reproducible benchmarking framework enabling fair and comparable evaluation of log parsers.
Existing log parsing approaches struggle to balance semantic understanding and computational efficiency: purely statistical methods lack semantic awareness, while full reliance on large language models (LLMs) incurs high latency and cost. This work proposes a dynamic routing mechanism that adaptively categorizes incoming logs into dense and sparse types, processing them respectively with efficient statistical pattern mining and lightweight LLM-based semantic reasoning. By minimizing unnecessary LLM invocations, the method maintains high parsing accuracy across 14 public datasets while achieving 7.9–18.6× speedup over pure LLM approaches and up to 1.5× faster execution than Drain. Furthermore, it reduces token consumption by 80.2%–94.1% and LLM call frequency by 86.4%–90.9%.
This work addresses the pervasive issue of redundant and isolated messages in system logs, which hinder downstream tasks such as model reasoning and anomaly detection. To tackle this challenge, the authors propose LogPurifier—the first task-agnostic log cleansing framework—that systematically purifies logs by extracting log templates and modeling their dependencies to accurately identify and remove messages irrelevant to system functional behavior. By doing so, LogPurifier enables effective log sanitization applicable across diverse analytical scenarios. Experimental results demonstrate that LogPurifier substantially improves both accuracy and efficiency in various downstream tasks, thereby validating its effectiveness and generalizability.
Existing log parsing methods rely heavily on handcrafted rules and statistical features, neglecting semantic information—leading to inaccurate template matching and poor generalization. To address this, we propose a structured template generation framework that synergistically integrates entropy-driven log clustering with large language model (LLM)-based chain-of-thought reasoning. Specifically, we introduce an information-entropy-guided automatic sampling strategy to replace manual rule design, and develop a semantics-aware chain-of-thought template merging mechanism that deeply embeds LLM inference capabilities into the entire template induction pipeline. Evaluated on multiple large-scale public benchmarks, our method achieves state-of-the-art performance, significantly improving parameter identification accuracy and template generalizability across diverse log formats. The implementation is publicly available.
Existing log analysis models are task-specific, rely heavily on domain-specific annotated data, exhibit poor generalization, and struggle with complex or unseen instructions. Method: We propose LogLM, an instruction-driven large language model for log analysis, which unifies diverse log tasks—including anomaly detection, parsing, and summarization—into a standardized instruction-response format. LogLM is adapted to the log domain via multi-task instruction tuning and log-specific instruction engineering. It accepts natural-language instructions and supports zero-shot cross-task transfer. Contribution/Results: Experiments demonstrate that LogLM outperforms all state-of-the-art methods across five core log analysis tasks. It exhibits strong generalization to complex instructions and previously unseen tasks. As a single unified model, LogLM replaces multiple specialized models, significantly improving deployment efficiency and task-agnostic capability.
This work addresses the scarcity of production-grade Security Operations Center (SOC) logs for research due to stringent privacy constraints, which has led prior studies to rely on synthetic or outdated data. To bridge this gap, the authors propose a methodology that, for the first time, transforms real-world financial-sector SIEM logs into reusable research artifacts while adhering to strict privacy boundaries. The approach preserves investigation-relevant structures through structured anonymization, mapping to the MITRE ATT&CK framework, deterministic validation, and large language model (LLM)-based behavioral compliance checks. The resulting artifact comprises 37 HIKARI challenges suitable for effective model training and demonstrates its utility by accurately identifying LLM policy violations across 200 SOCpilot incidents, thereby validating its balanced trade-off between privacy preservation and analytical fidelity.
Security logs exhibit diverse and semi-structured formats, making traditional parsing approaches heavily reliant on extensive engineering effort, while direct querying struggles to capture complex temporal patterns and cross-event semantics. This work proposes a natural language–to–log query code generation method that eliminates the need for custom parsers by leveraging lightweight, automatically extracted log format context to guide large language models in translating natural language security questions into executable query code. The approach requires only a single model invocation followed by deterministic execution. Evaluated across five log types and 133 security queries, the method reduces error rates by more than threefold compared to handcrafted scripts, demonstrating particularly significant improvements in critical tasks involving multi-line event correlations.
This work addresses the challenges of root cause diagnosis in large-scale microservice systems, where existing approaches are hindered by massive log volumes, limited LLM context windows, and insufficient semantic reasoning and interpretability. The authors propose a neuro-symbolic hybrid method that emulates Site Reliability Engineers’ manual troubleshooting process through a six-stage pipeline for log sampling, template clustering, and anomaly ranking, producing a concise evidence package for LLM-based root cause inference. This approach compresses raw logs by 1,000–7,000× while preserving critical failure signals and provides auditable log templates and statistical evidence, substantially enhancing interpretability and practicality. Evaluated on 11 real-world incidents, the method achieves an MRR of 0.790 and ranks the correct root cause within the top three candidates in over 90% of cases within one minute, earning strong endorsement from operations teams.
This work addresses the challenge of effectively analyzing massive, heterogeneous high-performance computing (HPC) logs, which hinders fault diagnosis and performance optimization. The authors propose a scalable log analysis workflow that uniquely integrates frequent pattern mining based on finite-state automata with job-level log correlation. By leveraging the Aho–Corasick automaton for efficient pattern storage and matching, and incorporating system hierarchy and message priority information, the approach enables automated detection and clustering of errors and anomalous events. Experiments on an exascale-class supercomputing system demonstrate that the method accurately identifies characteristic error sequences, reveals distinct failure patterns across different applications, and supports real-time, interpretable monitoring to enhance system resilience.
This work addresses the limitations of existing agent telemetry data, which, while useful for fault detection, struggles with precise root cause localization under insufficient evidence and lacks a reliable abstention mechanism. The authors introduce TelemetrySuffBench, a novel benchmark that decouples fault detection, source localization, and safe abstention for the first time. It employs multi-component delayed-binding trajectories, seven-factor telemetry masking, and ambiguous source pairs to systematically evaluate model capabilities. A unified protocol, candidate-set constraints, invalid-output statistics, and a frozen blind test set are implemented, along with an evidence-gating mechanism. Experiments reveal that source localization achieves Top-1 accuracy of 33.8%–97.2% with full telemetry but drops to ≤0.5% when using only metadata or standard views. Evidence gating reduces unjustified answers by 12.5–48.6 percentage points, highlighting a significant performance gap between detection and localization and underscoring the critical role of decision content and provenance information in reliable attribution.