design telemetry systems

Designs and implements end-to-end telemetry systems for software and infrastructure, including instrumentation (for example OpenTelemetry), collection and ingestion pipelines, correlation and aggregation of metrics/events/logs, and integration with logging and analytics backends. Builds real-time telemetry ingestion and monitoring pipelines and analyzes the collected telemetry to generate operational and product insights.

designtelemetrysystems

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$203K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of effectively integrating Internet of Things (IoT) data with business process event logs, which stems from their heterogeneous origins and differing levels of abstraction. Direct incorporation of raw IoT data often leads to excessive log complexity, thereby hindering process analysis. To overcome this, the paper proposes a structured fusion approach within the Object-Centric Event Log (OCEL) framework, introducing a mapping and integration mechanism that seamlessly embeds process-relevant IoT data into standard OCEL logs without requiring domain-specific schemas. The resulting IoT-enhanced event logs, generated by an extensible tool, are fully compatible with mainstream process mining tools. Empirical validation in real-world scenarios demonstrates the method’s effectiveness in enabling downstream process analysis and visualization.

business process analysisevent logsIoT integration

Interoperability From OpenTelemetry to Kieker: Demonstrated as Export from the Astronomy Shop

Oct 13, 2025
DG
David Georg Reichelt
🏛️ Lancaster University Leipzig | URZ Leipzig | Kiel University

Kieker’s observability capabilities are currently limited to a narrow set of languages (e.g., Java, C, Fortran, Python), hindering adoption in modern multi-language systems such as those using C# or JavaScript. To address this gap, we propose the first interoperability framework bridging OpenTelemetry and Kieker. Our approach introduces a distributed tracing data translation middleware that performs semantic mapping and format conversion from OpenTelemetry’s standardized telemetry protocol to Kieker’s event model. This enables unified ingestion of multi-language monitoring data into Kieker’s analysis stack, substantially extending its cross-language observability. Evaluation on the Astronomy Shop benchmark demonstrates accurate reconstruction and visualization of end-to-end call trees, validating both the completeness and practical utility of the transformation. Our work bridges a critical protocol compatibility gap in Kieker’s integration with contemporary observability ecosystems and provides a reusable technical pathway for retrofitting legacy analysis frameworks with emerging open standards.

Enabling call tree creation from OpenTelemetry instrumentationsTransforming OpenTelemetry tracing data into Kieker frameworkVisualizing trace data from OpenTelemetry demo applications

This work addresses the unreliability of developer productivity dashboards, which often stems from ad hoc scripts that introduce undetected silent data gaps, eroding organizational trust. To resolve this, we propose a robust ELT pipeline grounded in DAG-based orchestration and the Medallion architecture, decoupling data extraction from transformation to preserve the immutability of raw data. Our approach introduces a state-driven dependency scheduling mechanism and, for the first time, treats metric pipelines as production-grade distributed systems. We emphasize the critical role of immutable raw history in enabling reliable metric redefinition. This methodology significantly enhances data reliability and freshness while effectively eliminating silent failures, thereby restoring organizational confidence in DevOps metrics.

Data ReliabilityDeveloper ProductivityDORA Metrics

This work addresses the challenge scientists face in efficiently transforming raw sensor data streams into actionable insights across edge-cloud infrastructures, hindered by the need for cross-domain expertise to manage heterogeneous systems and emerging platforms such as DPUs, which impedes rapid prototyping. To overcome this barrier, the authors propose a novel paradigm that integrates pattern-based workflow engineering with AI-assisted development. Implemented on the FABRIC testbed using the Pegasus workflow system and exemplified by the Orcasound hydrophone workflow, this approach enables swift construction of applications for air quality, seismic, and soil moisture monitoring. The framework supports modular extensibility and edge deployment, substantially lowering the barrier for non-expert users to iteratively develop distributed applications. Empirical validation across multiple use cases demonstrates its effectiveness in enhancing development efficiency, accelerating prototyping cycles, and accumulating practical deployment experience.

cross-domain expertiseedge-to-cloud continuumheterogeneous infrastructure

PlantD: Performance, Latency ANalysis, and Testing for Data Pipelines -- An Open Source Measurement, Testing, and Simulation Framework

Apr 14, 2025
CB
Christopher Bogart
🏛️ Carnegie Mellon University | Honda Research Institute USA, Inc.

The proliferation of sensor devices (e.g., vehicle telematics) has led to rapidly growing data pipelines, making it difficult for enterprises to quantitatively predict infrastructure costs and performance for business teams—resulting in widespread over-provisioning. Method: We propose the “Data Pipeline Wind Tunnel” paradigm, integrating synthetic workload generation, multi-dimensional metric collection (latency, throughput, resource consumption), interactive visualization, and business-hypothesis-driven “what-if” modeling for annualized cost and SLA compliance. A reusable, open-source measurement harness is implemented to support systematic pipeline benchmarking. Contribution: This work establishes, for the first time, an interpretable mapping from engineering performance metrics to business decision parameters—including annualized infrastructure cost and SLA attainment rate. Evaluated across three real-world automotive data pipelines, the framework enables cross-functional collaboration and optimization, reducing infrastructure over-provisioning by up to 42% while maintaining SLA targets.

Closing gap between technical metrics and business cost implicationsForecasting cost and performance of data pipelines for device deploymentsSimulating pipeline performance under projected real-world loads

Latest Papers

What's happening recently
View more

This study addresses the lack of systematic evaluation of mainstream security logging standards in terms of their effectiveness for threat detection. The authors propose a scalable and reproducible assessment methodology based on an automated Security Exploit Telemetry Collection (SETC) framework, which reproduces 50 remote code execution vulnerabilities in containerized environments. Using this approach, they comparatively evaluate the telemetry completeness and attack detectability of widely adopted standards—including Common Information Model (CIM), Open Cybersecurity Schema Framework (OCSF), and Elastic Common Schema (ECS). The experiments quantitatively measure each standard’s detection efficacy, revealing significant disparities in coverage of critical attack indicators and identifying notable gaps. These findings provide empirical guidance for security practitioners in selecting appropriate logging standards to enhance threat detection capabilities.

cyber threat detectiondetection efficacylogging standards

This work addresses the challenge that existing experimental environments for distributed Cyber-Physical Systems (CPS) struggle to support reproducible, observable, and controllable integration of heterogeneous edge, fog, and cloud resources. To bridge this gap, the paper proposes a generic cloud continuum experimentation architecture grounded in the SLICES blueprint, featuring a two-layer reference model that decouples infrastructure from application logic. CPS workflows are structured along an edge–fog–cloud continuum, with deployment location, timing, and data provenance treated as core experimental dimensions. The architecture integrates virtualized and physical edge nodes, digital twin coordination, time-windowed control, and combined stream processing with cloud-side aggregation analytics, enabling multi-domain CPS applications to share programmable infrastructure and flexibly deploy and compare control and monitoring strategies. Validation through 40 systematic experiments across geographically distributed deployments—spanning renewable energy community management and AirWatch monitoring use cases—demonstrates the framework’s effectiveness and generality in hybrid physical-virtual settings.

Cloud ContinuumCyber-Physical SystemsDistributed Experimentation

Hot Scholars

GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
NS

Nicu Sebe

University of Trento
computer visionmultimedia
HX

Hui Xiong

Senior Scientist, Candela Corporation
Ultrafast dynamicsatomic molecular physicsfree electron laser
ZL

Ziwei Liu

Associate Professor, Nanyang Technological University
Computer VisionMachine LearningComputer Graphics
ZL

Zhenbo Luo

XiaoMi
Vision Language ModelComputer Vision