Optimized Log Parsing with Syntactic Modifications

๐Ÿ“… 2025-10-30
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenge of performance evaluation and optimization of log parsers by systematically comparing syntactic versus semantic approaches and single-stage versus two-stage architectures. We propose SynLog+, a lightweight template identification enhancement module designed as the second stage of a two-stage parsing framework; it jointly leverages syntactic analysis and semantic modeling to significantly improve accuracy with negligible runtime overhead. Experiments across diverse benchmarks demonstrate that SynLog+ boosts average accuracy by 236% for syntactic parsers and by 20% for semantic parsers, confirming its superior accuracyโ€“efficiency trade-off. Our core contributions are twofold: (1) the first generalizable, architecture-agnostic enhancement design for template identification within two-stage log parsing frameworks; and (2) a structured, reproducible benchmarking framework enabling fair and comparable evaluation of log parsers.

Technology Category

Natural Language Processing: Syntax โ€” Tagging, Chunking & ParsingSearch and Optimization: Learning to SearchMachine Learning: Statistical Relational/Logic Learning

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
๐Ÿ“ Abstract
Logs provide valuable insights into system runtime and assist in software development and maintenance. Log parsing, which converts semi-structured log data into structured log data, is often the first step in automated log analysis. Given the wide range of log parsers utilizing diverse techniques, it is essential to evaluate them to understand their characteristics and performance. In this paper, we conduct a comprehensive empirical study comparing syntax- and semantic-based log parsers, as well as single-phase and two-phase parsing architectures. Our experiments reveal that semantic-based methods perform better at identifying the correct templates and syntax-based log parsers are 10 to 1,000 times more efficient and provide better grouping accuracy although they fall short in accurate template identification. Moreover, two-phase architecture consistently improves accuracy compared to single-phase architecture. Based on the findings of this study, we propose SynLog+, a template identification module that acts as the second phase in a two-phase log parsing architecture. SynLog+ improves the parsing accuracy of syntax-based and semantic-based log parsers by 236% and 20% on average, respectively, with virtually no additional runtime cost.
Problem

Research questions and friction points this paper is trying to address.

Evaluating characteristics and performance of diverse log parsers
Comparing syntax-based versus semantic-based log parsing methods
Improving log parsing accuracy through two-phase architecture enhancements
Innovation

Methods, ideas, or system contributions that make the work stand out.

Two-phase architecture enhances log parsing accuracy
SynLog+ module improves syntax and semantic parsers
Minimal runtime cost for significant accuracy gains