Neural Representations, Natural Connections: What Transfers From Human Speech Foundation Models to Animal Vocalizations?

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether human speech foundation models can effectively transfer to animal bioacoustics tasks, examining the underlying cross-species transfer mechanisms and influencing factors. Leveraging multi-species datasets, the authors employ frozen encoders and controlled comparative experiments to systematically evaluate various pretraining strategies for animal call recognition and classification. The findings reveal that speech models exhibit notable cross-domain transfer potential, outperforming animal-specific models in certain scenarios. Furthermore, representation quality demonstrates significant layer dependence, whereas pretraining language coverage yields no systematic advantage. Crucially, human speech representations significantly outperform or match animal-pretrained representations across most experimental settings. These results underscore the viability of repurposing human speech foundation models for non-human bioacoustic analysis, offering new insights into cross-species acoustic representation learning.
📝 Abstract
Speech-pretrained models have shown promise in animal bioacoustics, but the factors governing their cross-species transfer remain poorly understood. We evaluate 15 frozen encoders spanning monolingual and multilingual speech, speaker verification, animal bioacoustics, and general audio on caller identification across four species and call-type classification across three. Differences between the strongest speech- and animal-pretrained representations range from $-0.027$ to $+0.119$ UAR; speech is significantly better in three of seven settings ($p<0.01$) and not significantly different in the remaining four. Controlled comparisons show no systematic advantage from increased language coverage or animal-domain pretraining, while frozen speaker-verification embeddings transfer poorly. The results also show that transfer is strongly layer-dependent: raw-waveform models peak early, whereas patch-spectrogram models peak deeper.
Problem

Research questions and friction points this paper is trying to address.

cross-species transfer
animal bioacoustics
speech foundation models
caller identification
call-type classification
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-species transfer
speech foundation models
animal bioacoustics
frozen encoders
layer-dependent transfer
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Tomás Arias-Vergara
Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Germany
C
Christopher Hauer
Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Germany; Faculty of Electrical Engineering, Media and Computer Science, OTH Amberg-Weiden, Germany
H
Héloïse Brotier
Faculty of Life Science, The Gonda Multidisciplinary Brain Research Center, Bar Ilan University, Israel
Elmar Nöth
Elmar Nöth
University of Erlangen-Nuremberg
speech
A
Andreas Maier
Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Germany
L
Lee Koren
Faculty of Life Science, The Gonda Multidisciplinary Brain Research Center, Bar Ilan University, Israel