tracing

Designs, implements, and evaluates instrumentation, logging and trace‑collection pipelines that record, correlate, store, and query execution or event traces across processes, threads, services, or system components. Analyzes those traces to diagnose faults, measure and attribute performance, understand control and data flow and provenance, and produce visualizations or summaries, addressing concerns such as sampling, aggregation, correlation, storage, and privacy.

tracing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Offloading tracing for real-time systems using a scalable cloud infrastructure

Jul 26, 2025
DJ
David Jannis Schmidt
🏛️ University of Applied Science | Schmidt Embedded Systems GmbH

Traditional local tracing tools face limitations in processing capability, real-time performance, and scalability. To address these challenges, this paper proposes a distributed, cloud-native tracing architecture leveraging microservices and edge computing. The architecture employs WebSocket for low-latency data ingestion and Apache Kafka for high-throughput, reliable transmission of embedded-system tracing data, enabling concurrent multi-session processing, lightweight streaming upload, and collaborative cloud-based analysis. Innovatively integrating edge-side preprocessing with cloud-side elastic scaling, it simultaneously supports rule-based validation during development and runtime fault diagnosis. Experimental results demonstrate that the system achieves linear throughput scaling with increasing load while maintaining stable per-session throughput. It effectively enables large-scale, long-term monitoring and collaborative debugging of intermittent faults. Consequently, the architecture significantly enhances observability and maintainability of real-time embedded systems.

Enabling long-term monitoring and collaborative error detectionOvercoming local desktop limitations in large-scale trace analysisScalable cloud-based tracing for real-time embedded systems

Trace Sampling 2.0: Code Knowledge Enhanced Span-level Sampling for Distributed Tracing

Sep 17, 2025
YW
Yulun Wu
🏛️ The Chinese University of Hong Kong

Distributed tracing data volume has surged, imposing prohibitive storage overhead; conventional trace-level sampling—e.g., retaining only anomalous traces—often discards critical diagnostic information such as normal execution paths. To address this, we propose Trace Sampling 2.0: the first approach to integrate code-knowledge-enhanced static analysis into distributed tracing, enabling span-level fine-grained sampling. We design an execution logic modeling mechanism and a structural consistency preservation scheme to compress traces while fully retaining call topology and key causal paths. Evaluated on two open-source microservice systems, our method achieves an 81.2% trace compression ratio, 98.1% recall for anomalous spans, and an average 8.3 percentage-point improvement in root-cause localization accuracy—demonstrating significant gains in both storage efficiency and diagnostic effectiveness.

Maintaining trace structural integrity while sampling at span levelPreserving valuable trace information discarded by current samplingReducing storage burden from high-volume distributed tracing

Tracing and Metrics Design Patterns for Monitoring Cloud-native Applications

Oct 03, 2025
CA
Carlos Albuquerque
🏛️ INESC TEC | Faculty of Engineering | University of Porto

Cloud-native systems face fragmented observability and challenging root-cause analysis due to their distributed, highly dynamic architectures. To address this, this paper proposes a reusable, observability-oriented design pattern system comprising three core categories: distributed tracing, application-level metric modeling, and infrastructure-level metric collection. Unlike ad-hoc toolchain integrations, our work is the first to systematically abstract industrial practices into structured, composable design patterns that holistically guide end-to-end monitoring architecture. Implemented and validated using mainstream frameworks—including OpenTelemetry and Prometheus—the approach significantly improves latency attribution accuracy, resource utilization assessment efficiency, and anomaly detection timeliness. Empirical evaluation across multiple microservice deployments demonstrates a 42% reduction in mean time to identify failures.

Improving visibility into distributed request flows for latency analysisMonitoring infrastructure environment for resource utilization and scalabilityProviding structured application instrumentation for real-time monitoring

UniSage: A Unified and Post-Analysis-Aware Sampling for Microservices

Sep 30, 2025
ZZ
Zhouruixing Zhu
🏛️ The Chinese University of Hong Kong, Shenzhen | The Chinese University of Hong Kong

In modern distributed systems, massive trace and log data incur prohibitive storage overhead and impede fault diagnosis. Existing presampling methods often discard failure-relevant signals, compromising diagnostic transparency. This paper proposes UniSage—the first unified trace and log sampling framework tailored for microservices—adopting a *post-analysis–aware* paradigm: lightweight multimodal anomaly detection and root cause analysis (RCA) are first executed on the full data stream to generate service-level diagnostic insights that guide sampling decisions. UniSage innovatively integrates two complementary pillars: *analysis-guided sampling*, prioritizing critical failure signals, and *edge-case preservation*, ensuring rare but potentially diagnostic behaviors are retained. Experiments demonstrate that at a 2.5% sampling rate, UniSage captures 56.5% of critical traces and 96.25% of relevant logs, improves RCA accuracy@1 by 42.45%, and processes 10 minutes of data in under 5 seconds—substantially outperforming state-of-the-art approaches.

Improving accuracy of root cause analysis in distributed systemsPreventing loss of failure-related information in sampling approachesReducing storage overhead from growing microservices traces and logs

Applying Process Mining on Scientific Workflows: a Case Study

Jul 06, 2023
ZS
Zahra Sadeghibogar
🏛️ RWTH Aachen University

SLURM logs in HPC scientific workflows lack explicit case identifiers, hindering direct application of process mining. Method: This paper proposes an automatic job-correlation method based on implicit job dependency modeling—parsing SLURM logs and jointly leveraging spatiotemporal job feature matching and graph-structured modeling to achieve end-to-end clustering of unannotated jobs. Contribution/Results: We introduce the first systematic preprocessing framework for process mining on HPC logs, integrating algorithms such as Heuristics Miner to support process discovery and bottleneck diagnosis. Evaluated on real-world HPC cluster logs, our approach significantly improves workflow traceability, accurately identifies I/O- and scheduler-related performance bottlenecks, and enables high-fidelity reconstruction of end-to-end process models.

Correlate jobs with explicit or implicit dependencies.Document workflows and identify performance bottlenecks.Extract case IDs from SLURM-based HPC logs.

Latest Papers

What's happening recently
View more

This work addresses the challenge of effectively analyzing massive, heterogeneous high-performance computing (HPC) logs, which hinders fault diagnosis and performance optimization. The authors propose a scalable log analysis workflow that uniquely integrates frequent pattern mining based on finite-state automata with job-level log correlation. By leveraging the Aho–Corasick automaton for efficient pattern storage and matching, and incorporating system hierarchy and message priority information, the approach enables automated detection and clustering of errors and anomalous events. Experiments on an exascale-class supercomputing system demonstrate that the method accurately identifies characteristic error sequences, reveals distinct failure patterns across different applications, and supports real-time, interpretable monitoring to enhance system resilience.

anomaly detectionHPC logslog analysis

Existing debugging tools excel at verifying hypotheses but struggle to support hypothesis generation, as programmers must manually reconstruct the program’s state evolution. This work proposes a novel debugging paradigm centered on complete execution traces, leveraging program tracing techniques to record and temporally visualize the actual code paths executed, rather than relying on the static structure of the source code. By presenting runtime behavior in a chronological and contextualized manner, this approach significantly enhances the comprehensibility of program execution, thereby facilitating more efficient hypothesis generation during debugging. We implement a prototype system and conduct preliminary experiments that demonstrate its effectiveness in improving program understanding efficiency, while also uncovering key challenges and promising directions for future research.

debuggingexecution tracehypothesis generation

Modern OLTP systems often suffer from frequent schema changes, missing primary/foreign keys, and fragmented execution traces, rendering traditional approaches—reliant on fixed schemas and manual modeling—costly and error-prone. This work proposes a fully automated pipeline that operates without predefined schemas by identifying quasi-key and timestamp columns, discovering inter-table relationships through statistical signals, and assembling and ordering events accordingly. To capture long-range dependencies across system events, the method incorporates a Temporal Convolutional Network (TCN). By eliminating dependence on ER diagrams, domain-specific templates, and stable schemas, the approach enables generalizable and scalable reconstruction of execution traces in dynamic information systems. Experimental results on TPC-H/E, synthetic, and real-world industrial datasets demonstrate 85% accuracy in event prediction and recovery of approximately 82% of true predecessor relationships, yielding high-fidelity process traces.

execution behaviorOLTPprocess trace construction

This work proposes an automated log aggregation and analysis framework based on large language models to address the growing challenge of log analysis in increasingly complex systems, where engineers traditionally rely on domain expertise to manually craft intricate LogQL queries. The framework enables end-to-end generation of LogQL queries from natural language instructions by integrating a hierarchical log knowledge base, natural language understanding, knowledge retrieval, and tool invocation mechanisms. Evaluated on four real-world log datasets, the approach achieves an average accuracy of 76.8%, significantly outperforming existing baselines and demonstrating its effectiveness and practicality for log analysis tasks.

DSL queryfault diagnosislog aggregation