log and telemetry correlation

Designs and implements systems and pipelines that join, align, and analyze log records and telemetry streams to identify related events, causal chains, and enriched context across sources. This includes building timestamp alignment and ID-mapping, defining correlation rules or algorithms, aggregations and alerting, and performing correlation analyses to reconstruct incident timelines and attribute anomalies.

logandtelemetrycorrelation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenges of root cause diagnosis in large-scale microservice systems, where existing approaches are hindered by massive log volumes, limited LLM context windows, and insufficient semantic reasoning and interpretability. The authors propose a neuro-symbolic hybrid method that emulates Site Reliability Engineers’ manual troubleshooting process through a six-stage pipeline for log sampling, template clustering, and anomaly ranking, producing a concise evidence package for LLM-based root cause inference. This approach compresses raw logs by 1,000–7,000× while preserving critical failure signals and provides auditable log templates and statistical evidence, substantially enhancing interpretability and practicality. Evaluated on 11 real-world incidents, the method achieves an MRR of 0.790 and ranks the correct root cause within the top three candidates in over 90% of cases within one minute, earning strong endorsement from operations teams.

incident diagnosislarge-scale systemslog analysis

Analyzing Logs of Large-Scale Software Systems using Time Curves Visualization

Nov 08, 2024
DB
Dmytro Borysenkov
🏛️ Dynatrace | Johannes Kepler University Linz

Analyzing massive, heterogeneous, and scale-free system logs remains challenging due to their volume, diversity, and lack of standardized structure. Method: This paper proposes Time Curves, a novel log analysis pipeline integrating log clustering, event detection, LLM-driven summarization, multidimensional scaling (MDS), and Time Curves visualization. It introduces a semi-metric distance function tailored for log events and is the first to apply Time Curves to log analysis—enabling concurrent, overlaid projections for temporal trend identification and fine-grained anomaly detection. Contribution/Results: The method requires no prior knowledge, automatically extracts dominant events from multi-source logs, and achieves joint semantic-temporal modeling with full interpretability. Evaluated on distributed systems, it simultaneously reveals global behavioral patterns and localized anomalies, significantly reducing time-to-diagnosis for faults, performance bottlenecks, and security threats.

Combining clustering and LLM for summarizationDetecting trends and outliers using Time CurvesEfficient log analysis for large-scale systems

Process Mining on Distributed Data Sources

Jun 03, 2025
MW
Maximilian Weisenseel
🏛️ Dresden University of Technology | Kiel University | Humboldt-Universität zu Berlin | Hamburg University of Technology | Utrecht University | University of Bayreuth

Traditional process mining relies on centralized, discrete event logs, rendering it inadequate for real-time, fine-grained, and heterogeneous event streams generated by distributed sensors in logistics, healthcare, and smart cities. Methodologically, this project advances three paradigm shifts: from offline to online, from centralized to distributed, and from log-based to sensor-driven process mining. It introduces the first distributed process intelligence framework tailored for continuous event streams, structured around a six-domain research agenda spanning infrastructure, data, and human-centered dimensions. The approach integrates formal modeling, distributed stream processing, privacy-preserving computation, and empirical evaluation, adhering to algorithm engineering principles for cross-layer co-design. Key contributions include a privacy-aware, scalable, and user-centric theoretical foundation and technology roadmap—enabling a new paradigm of responsive, decentralized process intelligence.

Adapting traditional techniques for online, decentralized sensor data analysisDeveloping scalable, privacy-aware methods for distributed environmentsExtending process mining to handle distributed, heterogeneous event streams

Existing evaluation approaches for streaming process mining algorithms predominantly rely on static logs or synthetic event streams, which fail to capture the complexity of real-world event streams in IoT environments—such as out-of-order events, concurrency, incomplete cases, and concept drift. This work addresses this gap by introducing, for the first time, a feature framework from data stream research into streaming process mining. It proposes an intent-oriented event stream generation methodology, extends the conceptual model of event streams, and implements a prototype tool, Stream of Intent. This tool enables customizable configuration of key stream characteristics reflective of real-world scenarios, facilitating the generation of controlled, reproducible, and realistically complex event streams. Consequently, it significantly enhances the relevance and adaptability of algorithm evaluation and development in streaming process mining.

BenchmarkingConcept DriftEvent Streams

Stochastic Alignments: Matching an Observed Trace to Stochastic Process Models

Jul 08, 2025
TL
Tian Li
🏛️ RWTH Aachen University | The University of Melbourne | Fraunhofer

Existing alignment-based conformance checking methods prioritize model paths with minimal edit distance to observed traces, neglecting their probabilistic plausibility—leading to low-probability, high-bias explanations. This paper proposes a path-likelihood-oriented alignment optimization framework that jointly minimizes edit distance and maximizes path probability within a stochastic process model, thereby identifying high-likelihood, low-deviation alignments. The method integrates stochastic process modeling, probability-aware heuristic search, and efficient optimization techniques. An open-source implementation demonstrates substantial improvements in alignment reasonableness and diagnostic utility across multiple business process scenarios. By unifying statistical rigor with interpretability, the approach establishes a novel paradigm for process deviation analysis.

Improving alignment-based conformance checking techniquesMatching observed traces to likely stochastic model pathsOptimizing low edit distance and high path likelihood

Latest Papers

What's happening recently
View more

This work addresses the challenge of effectively analyzing massive, heterogeneous high-performance computing (HPC) logs, which hinders fault diagnosis and performance optimization. The authors propose a scalable log analysis workflow that uniquely integrates frequent pattern mining based on finite-state automata with job-level log correlation. By leveraging the Aho–Corasick automaton for efficient pattern storage and matching, and incorporating system hierarchy and message priority information, the approach enables automated detection and clustering of errors and anomalous events. Experiments on an exascale-class supercomputing system demonstrate that the method accurately identifies characteristic error sequences, reveals distinct failure patterns across different applications, and supports real-time, interpretable monitoring to enhance system resilience.

anomaly detectionHPC logslog analysis

This work addresses the pervasive issue of redundant and isolated messages in system logs, which hinder downstream tasks such as model reasoning and anomaly detection. To tackle this challenge, the authors propose LogPurifier—the first task-agnostic log cleansing framework—that systematically purifies logs by extracting log templates and modeling their dependencies to accurately identify and remove messages irrelevant to system functional behavior. By doing so, LogPurifier enables effective log sanitization applicable across diverse analytical scenarios. Experimental results demonstrate that LogPurifier substantially improves both accuracy and efficiency in various downstream tasks, thereby validating its effectiveness and generalizability.

downstream tasksirrelevant messageslog analysis

This work addresses a critical limitation in existing log-based anomaly detection methods, which typically reduce logs to flat sequences of templates and thereby overlook the implicit multi-relational execution structures among events. To overcome this, the authors propose a novel approach that explicitly reconstructs an underlying state machine from raw logs and formulates a multi-table relational schema encompassing traces, events, states, transitions, and parameters. This schema guides synthetic data generation, preserving rare yet legitimate behaviors while adhering to structural, temporal, and procedural constraints. By uniquely integrating state machine discovery with multi-relational synthetic data generation, the method substantially enhances both the robustness and interpretability of anomaly detection. Experimental results demonstrate that the generated data outperforms baseline approaches in constraint satisfaction, distributional similarity, and workflow fidelity, leading to improved performance in detecting anomalies and defects on real-world datasets.

execution structurelog anomaly detectionrelational schema

This work addresses the challenge of distinguishing genuine causal relationships from mere correlations or temporal patterns in manufacturing system alarm logs. To this end, it proposes the LMT framework, which, for the first time, integrates semantic causal priors extracted by large language models with timestamp-based Poisson process likelihoods within a Bayesian causal discovery framework. By jointly modeling textual semantics and temporal statistical evidence, the method infers causal graphs that are both interpretable and well-supported by data. Extensive experiments demonstrate its superior performance across diverse simulated scenarios, particularly outperforming text-only or time-series-only baselines in low-data regimes with sparse alarm events.

Bayesian frameworkcausal discoveryevent causality

This work proposes an automated log aggregation and analysis framework based on large language models to address the growing challenge of log analysis in increasingly complex systems, where engineers traditionally rely on domain expertise to manually craft intricate LogQL queries. The framework enables end-to-end generation of LogQL queries from natural language instructions by integrating a hierarchical log knowledge base, natural language understanding, knowledge retrieval, and tool invocation mechanisms. Evaluated on four real-world log datasets, the approach achieves an average accuracy of 76.8%, significantly outperforming existing baselines and demonstrating its effectiveness and practicality for log analysis tasks.

DSL queryfault diagnosislog aggregation

Hot Scholars

XL

Xiapu Luo

The Hong Kong Polytechnic University
Mobile SecuritySmart ContractsNetwork SecurityBlockchain
YW

Yin Wu

Karlsruher Institut für Technologie
Autonomous DrivingADASScenario ExtractionAnomaly Detection
AS

Asaf Shabtai

Software and Information Systems Engineering, Telekom Innovation Labs, Ben Gurion University
Computer and network securitymachine learning
YV

Yash Vekaria

PhD Researcher, University of California at Davis
PrivacySecurityInternet MeasurementsLLMs
JL

Jiawei Liu

Wuhan University
Information RetrievalContent SecurityDocument Intelligence