Score
Design and build systems that analyze execution logs and other runtime traces to assign a labeled cause to individual failures, using supervised classifiers, ordered (priority) rule matching, or hybrid approaches; this includes log preprocessing, feature extraction from structured and unstructured text, model or rule construction, and evaluation of label quality (e.g., macro‑F1). Integrate the classifier or rule outputs into operational workflows (for example, attaching causes to completion notifications) and support reproducible diagnosis of observed failures.
This work addresses the pervasive issue of redundant and isolated messages in system logs, which hinder downstream tasks such as model reasoning and anomaly detection. To tackle this challenge, the authors propose LogPurifier—the first task-agnostic log cleansing framework—that systematically purifies logs by extracting log templates and modeling their dependencies to accurately identify and remove messages irrelevant to system functional behavior. By doing so, LogPurifier enables effective log sanitization applicable across diverse analytical scenarios. Experimental results demonstrate that LogPurifier substantially improves both accuracy and efficiency in various downstream tasks, thereby validating its effectiveness and generalizability.
Intermittent failures in continuous integration (CI) pipelines are notoriously difficult to diagnose, leading to wasted resources and reduced development efficiency. This work proposes FlaXifyer, a few-shot learning approach that integrates the interpretable AI technique LogSift to fine-tune pretrained language models on pipeline logs using only 12 labeled examples per failure class. The method simultaneously predicts failure categories and pinpoints critical log entries indicative of root causes. Evaluated on 2,458 real-world CI failures, FlaXifyer achieves a Macro F1 score of 84.3% and a Top-2 accuracy of 92.0%, reducing the required log inspection effort by 74.4%. Furthermore, it successfully identifies the underlying fault in 87% of cases, demonstrating its effectiveness in accelerating failure diagnosis with minimal labeled data.
This work addresses the challenges of efficiently analyzing large-scale, dynamically evolving semi-structured logs under conditions of label scarcity and distribution shift, which hinder system reliability and AIOps advancement. It presents the first unified task taxonomy for log analysis driven by large language models (LLMs), offering a systematic survey of their application across the full log analysis pipeline—including log generation, parsing, anomaly detection, and root cause analysis. Through structured analysis of 145 studies, the paper identifies five core design paradigms: prompt engineering, retrieval augmentation, fine-tuning, agent collaboration, and result verification. It further synthesizes the state of research, datasets, and evaluation practices across seven key tasks, while highlighting critical challenges in robustness, trustworthiness, and reproducibility, thereby providing a comprehensive roadmap for reliable LLM-based log intelligence.
Existing log analysis models are task-specific, rely heavily on domain-specific annotated data, exhibit poor generalization, and struggle with complex or unseen instructions. Method: We propose LogLM, an instruction-driven large language model for log analysis, which unifies diverse log tasks—including anomaly detection, parsing, and summarization—into a standardized instruction-response format. LogLM is adapted to the log domain via multi-task instruction tuning and log-specific instruction engineering. It accepts natural-language instructions and supports zero-shot cross-task transfer. Contribution/Results: Experiments demonstrate that LogLM outperforms all state-of-the-art methods across five core log analysis tasks. It exhibits strong generalization to complex instructions and previously unseen tasks. As a single unified model, LogLM replaces multiple specialized models, significantly improving deployment efficiency and task-agnostic capability.
Low-quality log statements—such as ambiguous or misleading ones—obscure actual program behavior and impede software maintenance. Prior work primarily focuses on detecting single log defects and relies on manual fixes. This paper proposes LogFixer, the first automated two-stage framework targeting four real-world log defects: detection and repair. In the offline stage, a lightweight similarity classifier is trained on synthetically defective logs; in the online stage, problematic logs are identified via joint modeling of static textual features and dynamic variable contexts, and semantically appropriate repairs are recommended using large language models (LLMs). LogFixer innovatively integrates a lightweight classifier with LLMs in a synergistic paradigm, ensuring robust detection while enhancing repair validity. Evaluation shows an F1-score of 0.625; adoption rates of static and dynamic repair suggestions improve by 48.12% and 24.90%, respectively; repair suggestion adoption reaches 61.49% on unseen projects; and 40 fixes submitted to GitHub have yielded 25 merged confirmations.
This work addresses the challenge of effectively analyzing massive, heterogeneous high-performance computing (HPC) logs, which hinders fault diagnosis and performance optimization. The authors propose a scalable log analysis workflow that uniquely integrates frequent pattern mining based on finite-state automata with job-level log correlation. By leveraging the Aho–Corasick automaton for efficient pattern storage and matching, and incorporating system hierarchy and message priority information, the approach enables automated detection and clustering of errors and anomalous events. Experiments on an exascale-class supercomputing system demonstrate that the method accurately identifies characteristic error sequences, reveals distinct failure patterns across different applications, and supports real-time, interpretable monitoring to enhance system resilience.
This study addresses the challenge of fault isolation from raw logs due to the lack of prior knowledge by proposing LOGOS, an unsupervised system that introduces a zero-prior diagnostic paradigm. Requiring neither seed queries nor predefined boundaries, LOGOS mines entity-event co-occurrences and temporal precedence relations to compress massive log volumes into a compact precedence forest structure. Experimental results demonstrate that LOGOS achieves a median processing time of only 4.5 minutes while eliminating 99.8% of noise and attaining a recall of 0.76. Furthermore, it provides warnings up to 16 hours in advance and improves the accuracy of LLM-based root cause analysis to 80%–100%. These findings indicate that LOGOS significantly reduces the reliance on domain expertise for troubleshooting proprietary software systems.
本文针对DevOps部署中的自动故障恢复难题,提出了一种结合基于规则和机器学习的混合框架,通过实时监控、识别、分类并恢复故障,有效提高了系统的恢复效率与运行稳定性。
This work addresses the challenges of root cause diagnosis in large-scale microservice systems, where existing approaches are hindered by massive log volumes, limited LLM context windows, and insufficient semantic reasoning and interpretability. The authors propose a neuro-symbolic hybrid method that emulates Site Reliability Engineers’ manual troubleshooting process through a six-stage pipeline for log sampling, template clustering, and anomaly ranking, producing a concise evidence package for LLM-based root cause inference. This approach compresses raw logs by 1,000–7,000× while preserving critical failure signals and provides auditable log templates and statistical evidence, substantially enhancing interpretability and practicality. Evaluated on 11 real-world incidents, the method achieves an MRR of 0.790 and ranks the correct root cause within the top three candidates in over 90% of cases within one minute, earning strong endorsement from operations teams.
本文提出了一种基于实际因果关系框架的方法,用于生成过程监控预测的局部解释,解决了AI模型黑盒问题,并在多个真实数据集上进行了验证。