🤖 AI Summary
This study investigates whether human speech foundation models can effectively transfer to animal bioacoustics tasks, examining the underlying cross-species transfer mechanisms and influencing factors. Leveraging multi-species datasets, the authors employ frozen encoders and controlled comparative experiments to systematically evaluate various pretraining strategies for animal call recognition and classification. The findings reveal that speech models exhibit notable cross-domain transfer potential, outperforming animal-specific models in certain scenarios. Furthermore, representation quality demonstrates significant layer dependence, whereas pretraining language coverage yields no systematic advantage. Crucially, human speech representations significantly outperform or match animal-pretrained representations across most experimental settings. These results underscore the viability of repurposing human speech foundation models for non-human bioacoustic analysis, offering new insights into cross-species acoustic representation learning.
📝 Abstract
Speech-pretrained models have shown promise in animal bioacoustics, but the factors governing their cross-species transfer remain poorly understood. We evaluate 15 frozen encoders spanning monolingual and multilingual speech, speaker verification, animal bioacoustics, and general audio on caller identification across four species and call-type classification across three. Differences between the strongest speech- and animal-pretrained representations range from $-0.027$ to $+0.119$ UAR; speech is significantly better in three of seven settings ($p<0.01$) and not significantly different in the remaining four. Controlled comparisons show no systematic advantage from increased language coverage or animal-domain pretraining, while frozen speaker-verification embeddings transfer poorly. The results also show that transfer is strongly layer-dependent: raw-waveform models peak early, whereas patch-spectrogram models peak deeper.