Uncertainty-Aware Crossmodal Fusion for Classification of Animal Behavior

📅 2026-08-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of animal sound classification in uncontrolled environments, where performance is degraded by noise, reverberation, vocalization overlap, and sensor deterioration. Existing acoustic representations—raw waveforms and log-Mel spectrograms—each exhibit inherent limitations. To overcome this, the authors propose an Uncertainty-Aware Fusion (UAF) framework that employs a dual-stream network to process both representations separately and estimates their Gaussian uncertainties in an unsupervised manner. The framework dynamically fuses the streams based on their estimated confidence, enabling uncertainty-guided integration without requiring reliability labels—a first in this domain. Evaluated under strict cross-species and leave-one-individual-out settings, UAF significantly outperforms static concatenation baselines, achieving macro F1 scores of 39.7% and 71.5% on the SoundWel and DogBark datasets, respectively, corresponding to relative improvements of 15.7% and 20.4%, thereby validating the efficacy of uncertainty-driven fusion.
📝 Abstract
Artificial intelligence offers substantial potential for acoustic monitoring of animals, from welfare assessment in precision livestock farming to wildlife conservation and ecological research, where vocalizations can indicate health, stress, and social states earlier and at lower cost than manual observation. However, recordings in these settings are obtained under uncontrolled conditions, including environmental noise, reverberation, overlapping calls, and sensors that degrade without notice. As a consequence, automated classification of animal vocalizations remains challenging, and the two dominant acoustic representations show complementary limitations: raw waveforms preserve temporal microstructure but degrade under clipping and reverberation, while log-Mel spectrograms capture harmonic organization but lose phase information and are sensitive to broadband noise. To address these challenges, we propose Uncertainty-Aware Fusion (UAF), a dual-stream framework that estimates Gaussian uncertainty for each representation and fuses them via uncertainty weighting. This mechanism assigns greater weight to the more confident representation with no reliability labels required. In a cross-species, identity-based evaluation excluding all individuals seen during training, UAF (mean pooling) achieves 59.4\% accuracy / 39.7\% macro F1 on the 17-class SoundWel pig vocalization benchmark and 73.1\% accuracy / 71.5\% macro F1 on the 3-class DogBark dataset, outperforming static-concatenation fusion by 15.7\% and 20.4\% relative macro F1, respectively. Ablations over four temporal aggregation strategies show that uncertainty fusion, rather than the temporal characteristics of animal calls, is the primary driver of the performance gain.
Problem

Research questions and friction points this paper is trying to address.

animal vocalization classification
acoustic monitoring
uncertainty-aware fusion
crossmodal representation
uncontrolled recording conditions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uncertainty-Aware Fusion
crossmodal fusion
animal vocalization classification
Gaussian uncertainty estimation
dual-stream framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Ehsan Yaghoubi
Ehsan Yaghoubi
Unknown affiliation
computer visioncrossmodal learning
F
Florian Haselbeck
Weihenstephan-Triesdorf University of Applied Sciences, Department of Sustainable Agricultural and Energy Systems, Freising, Am Staudengarten 1, 85354, Germany