Score
Design, build, and operate monitoring and alerting systems and automated pipelines that collect, aggregate, and track metrics, logs, traces and state from production deployments (including cloud-provider metrics), providing dashboards, audits and pipeline tracking. Implement anomaly detection and escalation rules to generate timely alerts, monitor operational and functional performance at scale, and support incident response and post‑hoc auditing.
This work addresses the high false positive rates and limited adaptability of traditional rule-based or static anomaly detection methods in cloud security logs, which struggle to capture the dynamic evolution of organizational behavior. The authors propose a self-supervised graph neural network approach that constructs user–resource interaction graphs from AWS CloudTrail logs to generate dynamic anomaly scores for each event. This method continuously adapts to environmental changes without requiring manual rule updates or frequent model retraining. Evaluated across five real-world organizations, the approach reduces hourly alerts from thousands to approximately one while maintaining high detection efficacy, substantially alleviating analyst workload and demonstrating strong practicality and adaptability in dynamic cloud environments.
This work addresses the unreliability of developer productivity dashboards, which often stems from ad hoc scripts that introduce undetected silent data gaps, eroding organizational trust. To resolve this, we propose a robust ELT pipeline grounded in DAG-based orchestration and the Medallion architecture, decoupling data extraction from transformation to preserve the immutability of raw data. Our approach introduces a state-driven dependency scheduling mechanism and, for the first time, treats metric pipelines as production-grade distributed systems. We emphasize the critical role of immutable raw history in enabling reliable metric redefinition. This methodology significantly enhances data reliability and freshness while effectively eliminating silent failures, thereby restoring organizational confidence in DevOps metrics.
This work addresses the threat posed by AI development agents that may covertly compromise infrastructure security through privilege escalation, log obfuscation, or persistence mechanisms—risks exacerbated by most organizations’ inability to deploy sophisticated monitoring. The authors propose a training-free, structured monitoring approach for Infrastructure-as-Code (IaC) environments, leveraging differential analysis between information flow graphs (IFGs) and control/data flow graphs to enable high-precision, interpretable detection of security degradation. The method supports both synchronous blocking and asynchronous auditing modes: in synchronous mode, real-time rollback reduces the success rate of stealthy attacks from 74.4% to 0%; in asynchronous mode, untrained IFGs lower the false-negative rate from 11.6% to 3.5% without disrupting legitimate operations.
This study addresses the critical challenge of alert fatigue in large-scale cloud systems, where excessive alerts severely degrade operational efficiency. To tackle this issue, the authors propose a novel three-stage framework that integrates large language models with lightweight graph learning to span the full lifecycle of alert management—encompassing alert denoising, summary generation, and iterative rule optimization. The approach innovatively combines graph-structured modeling (including virtual noise nodes) with retrieval-augmented generation and introduces a multi-agent feedback mechanism to enable continuous evolution of alert rules. Evaluated on real-world industrial datasets, the method achieves a 94.8% alert reduction rate and 90.5% fault diagnosis accuracy, successfully refining 1,174 alert rules, of which 375 were adopted by the Site Reliability Engineering (SRE) team.
This work addresses the challenges of anomaly detection in large-scale cloud system monitoring, where telemetry logs exhibit high-dimensional sparsity, intermittent service activity, and complex component dependencies. To tackle these issues, the authors propose ClouDens, a novel approach that leverages operational context attributes to guide domain-aware feature partitioning and construct context-aware graphs. ClouDens integrates a spatiotemporal graph neural network for predictive anomaly detection and incorporates sparse data imputation to enhance coverage. Experimental evaluation on a real-world IBM cloud telemetry dataset demonstrates that ClouDens significantly outperforms baseline models such as GRU on count-based features and achieves earlier, more accurate, and broader anomaly detection according to the NAB benchmark.
This work addresses the growing complexity of CI/CD pipelines and the lack of structured analysis capabilities in existing tools for understanding their behavior, failures, and version evolution. The authors propose an innovative approach that uniquely integrates digital twin technology with BPMN-based modeling in DevOps contexts. By automatically parsing raw CI configurations and execution logs, the method constructs structured, high-level process models that enable pipeline visualization, failure traceability, and cross-version comparison. Evaluated across multiple open-source projects, the approach demonstrates effectiveness in monitoring, evolutionary analysis, and fault diagnosis, offering a modular and extensible foundational framework for the analysis and optimization of CI/CD pipelines.
This study addresses the deployment latency caused by delayed detection of newly pushed container images in continuous deployment workflows. To mitigate this issue, the authors construct an end-to-end continuous deployment pipeline based on FluxCD and, for the first time, integrate SyMon into a real-world FluxCD environment to enable runtime monitoring of logs from GitHub Actions, GitHub Container Registry, FluxCD, and Kubernetes applications. Experimental results demonstrate that FluxCD consistently detects new images within 10 minutes, although detection within 5 minutes is not always reliable. SyMon effectively enables near-real-time monitoring, thereby validating its feasibility and practicality in quantifying image detection latency and ensuring timely deployments.
In highly regulated domains such as finance, data quality control (QC) is often fragmented into isolated preprocessing steps, undermining end-to-end trustworthy AI pipelines. To address this, we propose the first AI-driven DataOps framework that embeds QC as a system-level core component. Our framework deeply integrates rule-based engines, statistical analysis, and custom AI-powered anomaly detection across the entire data lifecycle—from ingestion and transformation to model deployment—enabling dynamic remediation, policy-configurable workflows, and end-to-end auditability. Technically, it unifies data profiling, stream processing, cloud-native storage interfaces, and a proprietary AI detection module. Evaluated in a real-world financial production environment, the framework achieves significantly improved anomaly recall, reduces manual intervention by 42%, and ensures audit completeness and full data traceability under high-throughput conditions, fully satisfying regulatory compliance requirements.
This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).