Score
Designs, implements, and evaluates observability frameworks and instrumentation that collect and process telemetry (metrics, logs, traces, events) to enable inference about a system's internal state and behavior. Builds data pipelines, dashboards, alerts, and analysis methods to diagnose faults, measure performance, and support root‑cause and longitudinal investigation.
To address insufficient observability in software systems—leading to slow fault localization and low operational efficiency—this paper designs and implements a lightweight, full-stack observability framework. The framework integrates Java bytecode instrumentation with event stream collection to enable runtime call-chain tracing, performance diagnostics, and root-cause analysis. It introduces a novel dual-mode deployment architecture supporting both online services and on-premises deployment, and achieves cross-toolchain collaborative visualization via tight REST API integration with ExplorViz. Evaluated on the TeaStore benchmark, the system delivers millisecond-scale distributed tracing and real-time heatmap rendering, reduces end-to-end latency by 32%, and shortens mean time to fault identification to the minute level. These results significantly enhance observability and operational intelligence for microservice systems.
In cloud-native microservices, manual and fragmented observability configuration leads to slow fault localization, high resource overhead, and degraded system performance. This paper introduces the first continuous observability assurance methodology, shifting from experience-driven to experiment-driven design. Built upon the Observability eXperimentation (OXN) framework, our approach integrates A/B testing, metric-based feedback loops, and Infrastructure-as-Code (IaC)-enabled automation to dynamically optimize and quantitatively evaluate observability configurations. Evaluated in realistic microservice deployments, our method reduces mean time to detection by 42% on average, decreases sampling overhead by 31%, and—uniquely—enables quantitative validation of how specific observability configurations directly impact Service-Level Objective (SLO) compliance. By establishing a reproducible, iterative, and empirically grounded design paradigm, this work advances observability engineering from ad hoc practice to rigorous, data-driven discipline.
Cloud-native systems face fragmented observability and challenging root-cause analysis due to their distributed, highly dynamic architectures. To address this, this paper proposes a reusable, observability-oriented design pattern system comprising three core categories: distributed tracing, application-level metric modeling, and infrastructure-level metric collection. Unlike ad-hoc toolchain integrations, our work is the first to systematically abstract industrial practices into structured, composable design patterns that holistically guide end-to-end monitoring architecture. Implemented and validated using mainstream frameworks—including OpenTelemetry and Prometheus—the approach significantly improves latency attribution accuracy, resource utilization assessment efficiency, and anomaly detection timeliness. Empirical evaluation across multiple microservice deployments demonstrates a 42% reduction in mean time to identify failures.
In cloud-native systems, alert rules frequently suffer from false positives and false negatives due to the absence of design-phase validation, while existing tools lack systematic support for alert testing. To address this, we propose the “Alert-as-Experiment” paradigm—the first adaptation of the observability experimentation framework OXN to early-stage alert rule validation. Our approach enables closed-loop, development-time testing and continuous calibration of alert logic via simulated execution, synthetic observation data injection, and real-world scenario replay. It supports parameter tuning and repeatable verification of alert-triggering behavior, shifting alert engineering from empirical practice toward a testable, verifiable, and systematic discipline. Empirical evaluation demonstrates significant reductions in both false positive and false negative rates in production environments, alongside improved fault response latency and system maintainability.
This work addresses a critical limitation in existing observability tools for multi-agent systems, which merely log interactions without enabling real-time enforcement of governance policies, thereby allowing violations to be detected only post hoc. To overcome this, the paper proposes a closed-loop governance architecture that embeds governance attributes directly into the telemetry layer, introducing a Governance-aware Telemetry System (GTS). By integrating an OPA-compatible declarative rule engine, a Governance Execution Bus (GEB), and a trusted telemetry plane, the system achieves real-time policy evaluation and tiered intervention with sub-200-millisecond latency. Leveraging encrypted provenance and enriched telemetry data, GTS establishes a tight feedback loop between observation and action, significantly enhancing runtime compliance and security in multi-agent AI systems.
GPU nodes are central to modern HPC and AI workloads, yet many failures do not manifest as immediate hard faults. While some instabilities emerge gradually as weak thermal or efficiency drift, a significant class occurs abruptly with little or no numeric precursor. In these detachment-class failures, GPUs become unavailable at the driver or interconnect level and the dominant observable signal is structural, including disappearance of device metrics and degradation of monitoring payload integrity. This paper proposes an observability-aware early-warning framework that jointly models (i) utilization-aware thermal drift signatures in GPU telemetry and (ii) monitoring-pipeline degradation indicators such as scrape latency increase, sample loss, time-series gaps, and device-metric disappearance. The framework is evaluated on production telemetry from GPU nodes at GWDG, where GPU, node, monitoring, and scheduler signals can be correlated. Results show that detachment failures exhibit minimal numeric precursor and are primarily observable through structural telemetry collapse, while joint modeling increases early-warning lead time compared to GPU-only detection. The dataset used in this study is publicly available at https://doi.org/10.5281/zenodo.19052367.
This work addresses the critical security challenges facing industrial control systems (ICS) as they become increasingly interconnected, while acknowledging the limitations of conducting research and testing in real-world environments. The authors propose an open-source, modular, and extensible ICS penetration testing platform that uniquely integrates network scanning, protocol-level interactions with Modbus and OPC UA, and large language model (LLM)-assisted report generation within a lightweight web interface. This enables asset discovery, enumeration, and controlled read/write operations. Notably, the platform leverages an LLM to automatically produce structured technical reports mapped to the MITRE ATT&CK for ICS framework, along with executive summaries. Experimental validation on a synthetic Modbus server, a Factory I/O water treatment scenario, and a custom OPC UA production line model demonstrates the platform’s effectiveness and practical utility.