logging

Designs and implements instrumentation and infrastructure for generating, collecting, storing, and analyzing log data (structured or unstructured event records), including log schemas, logging levels, rotation and retention policies, reliable transport, indexing, and query/alerting pipelines. Builds tools and processes for parsing, aggregating, searching, and visualizing logs to support debugging, monitoring, and incident investigation.

logging

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$195K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes an automated log aggregation and analysis framework based on large language models to address the growing challenge of log analysis in increasingly complex systems, where engineers traditionally rely on domain expertise to manually craft intricate LogQL queries. The framework enables end-to-end generation of LogQL queries from natural language instructions by integrating a hierarchical log knowledge base, natural language understanding, knowledge retrieval, and tool invocation mechanisms. Evaluated on four real-world log datasets, the approach achieves an average accuracy of 76.8%, significantly outperforming existing baselines and demonstrating its effectiveness and practicality for log analysis tasks.

DSL queryfault diagnosislog aggregation

To address the infeasibility of manual analysis for large-scale IT system logs, this paper proposes a lightweight log analysis framework leveraging large language models (LLMs). The method introduces a CPU-efficient inference mechanism that significantly improves LLM throughput on resource-constrained hardware without compromising semantic understanding fidelity. It integrates log parsing, contextual modeling, and fault-oriented semantic reasoning to enable end-to-end automated diagnosis. Deployed in production, the system supports 70 software products and has processed over 2,000 incident tickets. Empirical evaluation demonstrates an average monthly reduction of more than 300 human labor hours compared to conventional approaches—equivalent to approximately USD 15,444 in cost savings. The framework thus advances practical, scalable, and cost-effective LLM-based log analytics for real-world operational environments.

Automating log analysis for massive IT system logsProcessing large log volumes efficiently using LLMs on CPUsReducing manual effort in IT issue diagnosis and support

LLM-based event log analysis techniques: A survey

Feb 02, 2025
SA
Siraaj Akhtar
🏛️ University of Huddersfield

Traditional security log analysis methods suffer from low efficiency, high false-positive rates, and poor interpretability. Method: This paper presents the first systematic meta-analysis of large language model (LLM)-driven log analysis, synthesizing insights from 127 state-of-the-art studies through bibliometric analysis, methodological comparison, and cross-modal representation evaluation. Contribution/Results: We propose the first holistic taxonomy framework for LLM-based log analysis; identify critical gaps—including insufficient log format robustness and lack of causal reasoning—and derive design principles for scalable, standardized evaluation benchmarks. We categorize six mainstream technical paradigms (e.g., fine-tuning, retrieval-augmented generation, in-context learning), distill four persistent bottlenecks, and outline seven concrete future research directions. Our work delivers a theoretical roadmap and practical guidelines for automated threat detection and interpretable log auditing.

Computer Activity LogsEfficient AnalysisLarge Language Models

Optimized Log Parsing with Syntactic Modifications

Oct 30, 2025
NE
Nafid Enan
🏛️ York University

This study addresses the challenge of performance evaluation and optimization of log parsers by systematically comparing syntactic versus semantic approaches and single-stage versus two-stage architectures. We propose SynLog+, a lightweight template identification enhancement module designed as the second stage of a two-stage parsing framework; it jointly leverages syntactic analysis and semantic modeling to significantly improve accuracy with negligible runtime overhead. Experiments across diverse benchmarks demonstrate that SynLog+ boosts average accuracy by 236% for syntactic parsers and by 20% for semantic parsers, confirming its superior accuracy–efficiency trade-off. Our core contributions are twofold: (1) the first generalizable, architecture-agnostic enhancement design for template identification within two-stage log parsing frameworks; and (2) a structured, reproducible benchmarking framework enabling fair and comparable evaluation of log parsers.

Comparing syntax-based versus semantic-based log parsing methodsEvaluating characteristics and performance of diverse log parsersImproving log parsing accuracy through two-phase architecture enhancements

Latest Papers

What's happening recently
View more

Security logs exhibit diverse and semi-structured formats, making traditional parsing approaches heavily reliant on extensive engineering effort, while direct querying struggles to capture complex temporal patterns and cross-event semantics. This work proposes a natural language–to–log query code generation method that eliminates the need for custom parsers by leveraging lightweight, automatically extracted log format context to guide large language models in translating natural language security questions into executable query code. The approach requires only a single model invocation followed by deterministic execution. Evaluated across five log types and 133 security queries, the method reduces error rates by more than threefold compared to handcrafted scripts, demonstrating particularly significant improvements in critical tasks involving multi-line event correlations.

log parsinglog queryingsecurity logs

Natural language log querying remains challenging due to the absence of structured schemas, hindering accurate SQL generation. This work proposes a novel approach that first parses raw logs into templated relational tables and then enriches both templates and parameter columns with interpretable semantics through dual-granularity semantic grounding. By integrating semantic search with constrained decoding in large language models, the method generates context-aware, executable SQL queries. The study introduces the first semantically grounded log schema and releases LogNLQ-Bench, the inaugural benchmark for natural language log querying featuring execution-based validation. Experimental results demonstrate that the proposed method significantly outperforms existing techniques on LogNLQ-Bench, particularly excelling in complex analytical queries.

executable schemalog parsingnatural-language log querying

This work addresses the challenge of effectively analyzing massive, heterogeneous high-performance computing (HPC) logs, which hinders fault diagnosis and performance optimization. The authors propose a scalable log analysis workflow that uniquely integrates frequent pattern mining based on finite-state automata with job-level log correlation. By leveraging the Aho–Corasick automaton for efficient pattern storage and matching, and incorporating system hierarchy and message priority information, the approach enables automated detection and clustering of errors and anomalous events. Experiments on an exascale-class supercomputing system demonstrate that the method accurately identifies characteristic error sequences, reveals distinct failure patterns across different applications, and supports real-time, interpretable monitoring to enhance system resilience.

anomaly detectionHPC logslog analysis

This work addresses the pervasive issue of redundant and isolated messages in system logs, which hinder downstream tasks such as model reasoning and anomaly detection. To tackle this challenge, the authors propose LogPurifier—the first task-agnostic log cleansing framework—that systematically purifies logs by extracting log templates and modeling their dependencies to accurately identify and remove messages irrelevant to system functional behavior. By doing so, LogPurifier enables effective log sanitization applicable across diverse analytical scenarios. Experimental results demonstrate that LogPurifier substantially improves both accuracy and efficiency in various downstream tasks, thereby validating its effectiveness and generalizability.

downstream tasksirrelevant messageslog analysis

This work addresses the challenges of root cause diagnosis in large-scale microservice systems, where existing approaches are hindered by massive log volumes, limited LLM context windows, and insufficient semantic reasoning and interpretability. The authors propose a neuro-symbolic hybrid method that emulates Site Reliability Engineers’ manual troubleshooting process through a six-stage pipeline for log sampling, template clustering, and anomaly ranking, producing a concise evidence package for LLM-based root cause inference. This approach compresses raw logs by 1,000–7,000× while preserving critical failure signals and provides auditable log templates and statistical evidence, substantially enhancing interpretability and practicality. Evaluated on 11 real-world incidents, the method achieves an MRR of 0.790 and ranks the correct root cause within the top three candidates in over 90% of cases within one minute, earning strong endorsement from operations teams.

incident diagnosislarge-scale systemslog analysis