๐ค AI Summary
This study addresses the scarcity of high-quality labeled data for audio classification in domain-specific scenarios such as domestic environments. To this end, the authors propose TriA Pipeline, the first large-scale automated audio annotation framework tailored for such settings, which efficiently generates the TriA dataset comprising 2,130 hours of audio across 431 event classes. Furthermore, they introduce a prior knowledgeโguided data filtering mechanism to construct a refined subset, TriA_GK. Experimental results on three household audio classification tasks demonstrate that models trained on TriA_GK achieve relative improvements of 3.97% in average accuracy and 3.35% in Macro-F1 score over baseline methods, highlighting the effectiveness of the proposed approach.
๐ Abstract
There are some datasets of varying scales for audio classification (AC) applied to different tasks. However, annotated data is limited for most scenarios, such as domestic environments. To address this challenge, we propose an $\textbf{A}$utomatic $\textbf{A}$udio $\textbf{A}$nnotation Pipeline--TriA Pipeline, which can efficiently convert audio from various scenarios into high-quality training data with audio event annotations. A TriA dataset was constructed with the TriA Pipeline, over 2130 hours of audio covering 431 audio classes. Furthermore, we partitioned a prior-knowledge-guided subset (TriA$_{\mathrm{GK}}$) from TriA and conduct comparative experiments on three domestic AC tasks. Comparing the result on manually annotated data only and that on manually annotated data combines TriA$_{\mathrm{GK}}$, TriA$_{\mathrm{GK}}$ could achieve average relative gains of 3.97% in accuracy and 3.35% in Macro-F1, validating the effectiveness of TriA$_{\mathrm{GK}}$ and the TriA Pipeline.