logging and tracing

Designs and implements logging and tracing infrastructures and pipelines that collect, parse, structure, index, and retain logs, traces, and related metrics (including audit logs and structured logging). Builds and operates log analysis, parsing, forensics, large-scale processing and storage, monitoring/alerting integrations, and tracing systems to support incident investigation, performance analysis, and compliance reporting.

loggingandtracing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.57
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$192K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address the infeasibility of manual analysis for large-scale IT system logs, this paper proposes a lightweight log analysis framework leveraging large language models (LLMs). The method introduces a CPU-efficient inference mechanism that significantly improves LLM throughput on resource-constrained hardware without compromising semantic understanding fidelity. It integrates log parsing, contextual modeling, and fault-oriented semantic reasoning to enable end-to-end automated diagnosis. Deployed in production, the system supports 70 software products and has processed over 2,000 incident tickets. Empirical evaluation demonstrates an average monthly reduction of more than 300 human labor hours compared to conventional approaches—equivalent to approximately USD 15,444 in cost savings. The framework thus advances practical, scalable, and cost-effective LLM-based log analytics for real-world operational environments.

Automating log analysis for massive IT system logsProcessing large log volumes efficiently using LLMs on CPUsReducing manual effort in IT issue diagnosis and support

Logging Requirement for Continuous Auditing of Responsible Machine Learning-based Applications

Aug 25, 2025
PL
Patrick Loic Foalem
🏛️ Polytechnique Montreal

Machine learning systems lack auditability concerning transparency, fairness, and accountability. Method: This paper introduces a novel approach that systematically embeds responsible AI metrics—such as bias, explainability, and decision provenance—into logging infrastructure. Unlike conventional operational logs, the proposed framework integrates software engineering logging practices with AI ethics assessment dimensions, yielding a structured log model enabling continuous monitoring, traceable verification, and dynamic compliance checking. Contribution/Results: It represents the first method to achieve deep synergy between AI governance metrics and logging infrastructure, bridging the audit gap between model behavior and ethical compliance. Empirical evaluation demonstrates significant improvements in verifiability during regulatory audits and stakeholder trust. The approach provides actionable, implementation-ready guidance for developers and toolchain designers to enhance algorithmic accountability.

Addressing transparency and accountability in ML decision-making systemsDeveloping logging practices for continuous auditing of responsible AIEnhancing compliance with ethical and legal requirements in ML

This work addresses the challenges of efficiently analyzing large-scale, dynamically evolving semi-structured logs under conditions of label scarcity and distribution shift, which hinder system reliability and AIOps advancement. It presents the first unified task taxonomy for log analysis driven by large language models (LLMs), offering a systematic survey of their application across the full log analysis pipeline—including log generation, parsing, anomaly detection, and root cause analysis. Through structured analysis of 145 studies, the paper identifies five core design paradigms: prompt engineering, retrieval augmentation, fine-tuning, agent collaboration, and result verification. It further synthesizes the state of research, datasets, and evaluation practices across seven key tasks, while highlighting critical challenges in robustness, trustworthiness, and reproducibility, thereby providing a comprehensive roadmap for reliable LLM-based log intelligence.

AIOpsdata driftlarge language models

This work addresses the lack of traceable and tamper-resistant transparency mechanisms in large language models (LLMs) deployed in high-stakes decision-making contexts, which undermines accountability. To bridge this gap, the paper introduces the first LLM lifecycle auditing framework that integrates technical provenance with governance records. It proposes a reference architecture enabling cross-organizational traceability and implements a lightweight, open-source Python-based auditing layer. By leveraging append-only logs, event emitters, structured metadata, and an auditor interface, the system seamlessly integrates into existing LLM workflows with minimal intrusiveness. This design ensures complete, tamper-evident traceability across critical stages—including training, deployment, and monitoring—thereby facilitating robust accountability and responsibility attribution throughout the model’s lifecycle.

accountabilityaudit trailsgovernance

This work addresses the pervasive issue of redundant and isolated messages in system logs, which hinder downstream tasks such as model reasoning and anomaly detection. To tackle this challenge, the authors propose LogPurifier—the first task-agnostic log cleansing framework—that systematically purifies logs by extracting log templates and modeling their dependencies to accurately identify and remove messages irrelevant to system functional behavior. By doing so, LogPurifier enables effective log sanitization applicable across diverse analytical scenarios. Experimental results demonstrate that LogPurifier substantially improves both accuracy and efficiency in various downstream tasks, thereby validating its effectiveness and generalizability.

downstream tasksirrelevant messageslog analysis

Latest Papers

What's happening recently
View more

This study addresses the longstanding disconnect between detection engineering and digital forensics, which has led to a gap between real-time alerts and post-incident analysis. To bridge this divide, the authors propose a unified detection-and-forensics methodology based on Velociraptor that triggers targeted evidence collection immediately upon detection events, thereby integrating monitoring and forensic workflows. The approach introduces an innovative four-stage framework that transforms forensic artifacts into reusable, testable detection rules, enabling efficient initial triage without requiring full disk imaging. By leveraging BaseVQL data sources—such as Prefetch, USN Journal, and WMI—it facilitates cross-artifact correlation and periodic analysis, allowing effective screening even in the absence of Windows Event Logs. This significantly reduces data acquisition volume while supporting continuous monitoring.

artefact analysisdetection engineeringdigital forensics

This work proposes an automated log aggregation and analysis framework based on large language models to address the growing challenge of log analysis in increasingly complex systems, where engineers traditionally rely on domain expertise to manually craft intricate LogQL queries. The framework enables end-to-end generation of LogQL queries from natural language instructions by integrating a hierarchical log knowledge base, natural language understanding, knowledge retrieval, and tool invocation mechanisms. Evaluated on four real-world log datasets, the approach achieves an average accuracy of 76.8%, significantly outperforming existing baselines and demonstrating its effectiveness and practicality for log analysis tasks.

DSL queryfault diagnosislog aggregation

This work addresses the challenge of effectively analyzing massive, heterogeneous high-performance computing (HPC) logs, which hinders fault diagnosis and performance optimization. The authors propose a scalable log analysis workflow that uniquely integrates frequent pattern mining based on finite-state automata with job-level log correlation. By leveraging the Aho–Corasick automaton for efficient pattern storage and matching, and incorporating system hierarchy and message priority information, the approach enables automated detection and clustering of errors and anomalous events. Experiments on an exascale-class supercomputing system demonstrate that the method accurately identifies characteristic error sequences, reveals distinct failure patterns across different applications, and supports real-time, interpretable monitoring to enhance system resilience.

anomaly detectionHPC logslog analysis

This work addresses the challenges of root cause diagnosis in large-scale microservice systems, where existing approaches are hindered by massive log volumes, limited LLM context windows, and insufficient semantic reasoning and interpretability. The authors propose a neuro-symbolic hybrid method that emulates Site Reliability Engineers’ manual troubleshooting process through a six-stage pipeline for log sampling, template clustering, and anomaly ranking, producing a concise evidence package for LLM-based root cause inference. This approach compresses raw logs by 1,000–7,000× while preserving critical failure signals and provides auditable log templates and statistical evidence, substantially enhancing interpretability and practicality. Evaluated on 11 real-world incidents, the method achieves an MRR of 0.790 and ranks the correct root cause within the top three candidates in over 90% of cases within one minute, earning strong endorsement from operations teams.

incident diagnosislarge-scale systemslog analysis

Hot Scholars

JR

Ji-Rong Wen

Gaoling School of Artificial Intelligence, Renmin University of China
Large Language ModelWeb SearchInformation RetrievalMachine Learning
RT

Radu Timofte

Humboldt Professor for AI and Computer Vision, University of Würzburg
Computer VisionMachine LearningAICompression
DI

Dmitry Ignatov

Associate Professor, MMCP Lab Head, Computer Science Faculty, Higher School of Economics
Data MiningMachine LearningFormal Concept AnalysisAI
ZT

Zhengzhong Tu

Texas A&M University, Google Research, University of Texas at Austin
Agentic AITrustworthy AIEmbodied AI
TL

Tongliang Liu

Director, Sydney AI Centre, University of Sydney & Mohamed bin Zayed University of AI
Machine LearningLearning with Noisy LabelsTrustworthy Machine Learning