Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset

📅 2024-11-13
🏛️ arXiv.org
📈 Citations: 0
Influential: 0
📄 PDF

career value

251K/year
🤖 AI Summary
Large-scale cloud systems lack realistic, high-dimensional, long-term benchmark datasets for anomaly detection, hindering rigorous evaluation and advancement of industrial-strength methods. Method: We introduce the first ultra-large-scale, production-derived anomaly detection dataset from IBM Cloud—comprising 39,000 time series across 117,000 timestamps—featuring multi-source cloud-native telemetry (e.g., CPU, memory, latency, distributed traces) and fine-grained annotations of complex, real-world anomaly patterns. Leveraging this dataset, we propose a hybrid anomaly modeling framework integrating unsupervised and weakly supervised learning, accompanied by reproducible baseline results and a systematic analysis of practical challenges. Contribution/Results: Our dataset fills a critical industry gap, while experiments demonstrate substantial improvements in model generalization to real cloud environments and efficiency in industrial-grade evaluation. This work establishes a foundational data resource and methodological framework for advancing reliability research in cloud systems.

Technology Category

Application Category

📝 Abstract
As Large-Scale Cloud Systems (LCS) become increasingly complex, effective anomaly detection is critical for ensuring system reliability and performance. However, there is a shortage of large-scale, real-world datasets available for benchmarking anomaly detection methods. To address this gap, we introduce a new high-dimensional dataset from IBM Cloud, collected over 4.5 months from the IBM Cloud Console. This dataset comprises 39,365 rows and 117,448 columns of telemetry data. Additionally, we demonstrate the application of machine learning models for anomaly detection and discuss the key challenges faced in this process. This study and the accompanying dataset provide a resource for researchers and practitioners in cloud system monitoring. It facilitates more efficient testing of anomaly detection methods in real-world data, helping to advance the development of robust solutions to maintain the health and performance of large-scale cloud infrastructures.
Problem

Research questions and friction points this paper is trying to address.

Large-scale Cloud Systems
Issue Identification
Data Set Limitations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cloud System Dataset
Machine Learning Analysis
Performance Optimization
🔎 Similar Papers
No similar papers found.
Mohammad Saiful Islam
Mohammad Saiful Islam
Google
Security
M
M. Rakha
Dept. of Computer Science, Toronto Metropolitan University, Toronto, Canada
W
William Pourmajidi
Dept. of Computer Science, Toronto Metropolitan University, Toronto, Canada
J
Janakan Sivaloganathan
Dept. of Computer Science, Toronto Metropolitan University, Toronto, Canada
J
John Steinbacher
Cloud Platform, IBM Canada Lab, Toronto
A
Andriy V. Miranskyy
Dept. of Computer Science, Toronto Metropolitan University, Toronto, Canada