Clustering algorithms for multivariate wind farm SCADA data filtering

📅 2026-07-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of contaminated SCADA data from wind turbines, which often includes anomalies, transients, and non-steady-state operating points, rendering traditional expert-driven manual filtering inefficient. To overcome this limitation, the authors propose an unsupervised filtering approach that leverages multivariate feature engineering combined with multiple clustering algorithms to automatically detect both explicit and implicit abnormal operating conditions. A key innovation lies in the design of robust evaluation metrics tailored for unlabeled SCADA data, moving beyond the conventional reliance on power curves alone. Experimental results demonstrate that the proposed method consistently outperforms manual filtering across most scenarios, preserving a higher proportion of valid operational data while requiring only minimal expert calibration for deployment.
📝 Abstract
During wind farm operation, Supervisory Control and Data Acquisition (SCADA) systems record numerous anomalies, transients, and specific operational modes, leading to large datasets. However, for a wide range of applications, only measurements corresponding to normal operation are required and, therefore, the SCADA data must be filtered. For this purpose, several methods have been proposed to automate and replace manual filtering conducted by experts via visual inspection of the data. In this paper, we compare the filtering accuracy of multiple clustering algorithms against manual filtering, introducing evaluation metrics that are suitable for unlabeled data and robust across potential applications. Based on the results, we provide recommendations for generalizing model calibration to different datasets and discuss potential use cases for each model. The models are applied to the SCADA data of three turbines of an existing offshore wind farm, using 10-minute statistics across multiple data channels. In addition to the anomalies and operational modes typically recorded, the dataset presents a large number of non-evident outliers due to several field tests. Overall, the results highlight the importance of extending the analysis beyond the power curve, both in feature selection and in the design of evaluation metrics. In most cases, cluster-based methods are able to detect both evident and subtle outliers, achieving higher accuracy than manual filtering. However, the accuracy and the amount of data retained vary considerably depending on the model, and expert involvement remains necessary, though to a reduced extent compared to manual filtering.
Problem

Research questions and friction points this paper is trying to address.

SCADA data filtering
wind farm
clustering algorithms
anomaly detection
normal operation
Innovation

Methods, ideas, or system contributions that make the work stand out.

clustering algorithms
SCADA data filtering
unsupervised evaluation metrics
multivariate wind farm data
outlier detection
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
N
Nicolò Italiano
Department of Wind and Energy Systems, Technical University of Denmark, Frederiksborgvej 399, Roskilde, 4000, Denmark
V
Vasilis Pettas
Department of Wind and Energy Systems, Technical University of Denmark, Frederiksborgvej 399, Roskilde, 4000, Denmark
T
Tuhfe Göçmen
Department of Wind and Energy Systems, Technical University of Denmark, Frederiksborgvej 399, Roskilde, 4000, Denmark
Nicolaos A. Cutululis
Nicolaos A. Cutululis
Professor @ Technical University of Denmark
Wind poweroffshore windintegrationcontrolHVDC