🤖 AI Summary
Traditional bioacoustic classification relies on fixed-rate resampling, resulting in high-frequency information loss and poor compatibility with heterogeneous data. This work proposes a sampling frequency-independent (SFI) frontend coupled with a Fourier Neural Operator (FNO) backbone that directly processes native-sampling-rate audio through continuous filterbanks and progressive temporal scale fusion, complemented by training-time sampling rate augmentation to enhance robustness. Evaluated on a dataset encompassing 84 classes across 60 distinct sampling rates, the proposed method achieves 90.6% accuracy, 92.1% balanced accuracy, and an 89.9% macro F1 score, significantly outperforming existing baseline models.
📝 Abstract
Conventional bioacoustic classification models rely on fixed-rate spectral representations, requiring recordings acquired at heterogeneous sampling rates to be resampled before analysis. We propose a Sampling-Frequency-Independent (SFI) frontend that processes each recording directly at its native sampling rate, coupled with a Fourier Neural Operator (FNO) backbone featuring progressive temporal-scale fusion. This framework avoids fixed-rate resampling and high-frequency information loss while producing fixed-size representations across sampling rates. Mild training-time sampling-rate (\textit{sr}) augmentation further improves robustness to unseen rate variations. Evaluated on a multi-taxa corpus comprising 84 classes and 60 sampling rates, the proposed SFI-FNO configuration outperforms fixed-rate and corpus-maximum-rate baselines, achieving .906 accuracy, .921 balanced accuracy, and a Macro-F1 score of .899.