🤖 AI Summary
Standardized algorithm evaluation protocols are lacking in audio-based device health monitoring, hindering reproducible research and industrial deployment. This paper introduces the first standardized machine learning evaluation framework for this task: it systematically extracts 127 acoustic features spanning time, frequency, and time–frequency domains; rigorously benchmarks 12 classifiers under a unified cross-validation protocol and nonparametric statistical testing (e.g., Wilcoxon signed-rank test) on both synthetic and real-world industrial datasets; and proposes a novel ensemble strategy achieving 94.2% accuracy and an F1-score of 0.942 across diverse scenarios—significantly outperforming the best individual classifier (+8–15 percentage points, *p* < 0.01). The framework is open-sourced with a comprehensive benchmarking protocol, providing a reproducible, statistically validated foundation for principled algorithm selection and practical deployment.
📝 Abstract
Audio-based equipment condition monitoring suffers from a lack of standardized methodologies for algorithm selection, hindering reproducible research. This paper addresses this gap by introducing a comprehensive framework for the systematic and statistically rigorous evaluation of machine learning models. Leveraging a rich 127-feature set across time, frequency, and time-frequency domains, our methodology is validated on both synthetic and real-world datasets. Results demonstrate that an ensemble method achieves superior performance (94.2% accuracy, 0.942 F1-score), with statistical testing confirming its significant outperformance of individual algorithms by 8-15%. Ultimately, this work provides a validated benchmarking protocol and practical guidelines for selecting robust monitoring solutions in industrial settings.