A Story About Cohesion and Separation: Label-Free Metric for Log Parser Evaluation

📅 2025-12-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing log parsers are evaluated primarily using manually annotated ground truths, leading to dataset stagnation, poor industrial applicability, and inconsistent evaluation outcomes across differing ground-truth versions. To address this, we propose PMSS—the first label-agnostic, template-level evaluation metric—integrating Medoid silhouette analysis with Levenshtein distance to jointly quantify clustering cohesion and template semantic fidelity. Evaluated on LogHub 2.0, PMSS exhibits highly significant positive correlation with mainstream metrics FGA and FTA (ρ > 0.58, p < 1e−8), with only a 2.1% performance gap; it scales nearly linearly and operates effectively in both unlabeled and multi-version ground-truth settings. This work establishes a novel, ground-truth-free paradigm for robust parser evaluation and provides an interpretable framework for parser selection.

Technology Category

Machine Learning: Evaluation and AnalysisNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsSearch and Optimization: Evaluation and Analysis

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methodsSecurity and Privacy: Large-scale security measurements
📝 Abstract
Log parsing converts log messages into structured event templates, allowing for automated log analysis and reducing manual inspection effort. To select the most compatible parser for a specific system, multiple evaluation metrics are commonly used for performance comparisons. However, existing evaluation metrics heavily rely on labeled log data, which limits prior studies to a fixed set of datasets and hinders parser evaluations and selections in the industry. Further, we discovered that different versions of ground-truth used in existing studies can lead to inconsistent performance conclusions. Motivated by these challenges, we propose a novel label-free template-level metric, PMSS (parser medoid silhouette score), to evaluate log parser performance. PMSS evaluates both parser grouping and template quality with medoid silhouette analysis and Levenshtein distance within a near-linear time complexity in general. To understand its relationship with label-based template-level metrics, FGA and FTA, we compared their evaluation outcomes for six log parsers on the standard corrected Loghub 2.0 dataset. Our results indicate that log parsers achieving the highest PMSS or FGA exhibit comparable performance, differing by only 2.1% on average in terms of the FGA score; the difference is 9.8% for FTA. PMSS is also significantly (p<1e-8) and positively correlated to both FGA and FTA: the Spearman's rho correlation coefficient of PMSS-FGA and PMSS-FTA are respectively 0.648 and 0.587, close to the coefficient between FGA and FTA (0.670). We further extended our discussion on how to interpret the conclusions from different metrics, identifying challenges in using PMSS, and provided guidelines on conducting parser selections with our metric. PMSS provides a valuable evaluation alternative when ground-truths are inconsistent or labels are unavailable.
Problem

Research questions and friction points this paper is trying to address.

Proposes a label-free metric for evaluating log parser performance.
Addresses reliance on labeled data and inconsistent ground-truth versions.
Enables parser selection without labeled datasets in industrial settings.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proposes PMSS, a label-free metric for log parser evaluation
Uses medoid silhouette analysis and Levenshtein distance for assessment
Operates with near-linear time complexity for efficient performance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Qiaolin Qin
Qiaolin Qin
Polytechnique Montreal
Data EngineeringLog AnalysisAI Fairness
J
Jianchen Zhao
University of Waterloo, Waterloo, Canada
H
Heng Li
Polytechnique Montreal, Montreal, Canada
W
Weiyi Shang
University of Waterloo, Waterloo, Canada
E
Ettore Merlo
Polytechnique Montreal, Montreal, Canada