🤖 AI Summary
This work addresses the underexplored audio-only scenario by jointly tackling sound event localization (DOA estimation), detection, and distance estimation for the first time. We propose a ResNet-Conformer hybrid architecture built upon EINV2, which fuses log-Mel spectrograms with binaural intensity vectors. To unify spatio-temporal and spatial modeling, we incorporate weight-sharing mechanisms and self-attention modules. A multi-task learning framework—augmented with task-specific data augmentation—enables end-to-end joint prediction of direction and distance. Evaluated on the Development Dataset, our method achieves 40.2% F-score, 17.7° DOA error, and 0.32 relative distance error—marking substantial improvements in 3D acoustic event perception from audio alone. This work establishes a novel paradigm for vision-free sound source understanding, advancing robust auditory scene analysis without visual cues.
📝 Abstract
This technical report outlines our approach to Task 3A of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024, focusing on Sound Event Localization and Detection (SELD). SELD provides valuable insights by estimating sound event localization and detection, aiding in various machine cognition tasks such as environmental inference, navigation, and other sound localization-related applications. This year's challenge evaluates models using either audio-only (Track A) or audiovisual (Track B) inputs on annotated recordings of real sound scenes. A notable change this year is the introduction of distance estimation, with evaluation metrics adjusted accordingly for a comprehensive assessment. Our submission is for Task A of the Challenge, which focuses on the audio-only track. Our approach utilizes log-mel spectrograms, intensity vectors, and employs multiple data augmentations. We proposed an EINV2-based [1] network architecture, achieving improved results: an F-score of 40.2%, Angular Error (DOA) of 17.7 degrees, and Relative Distance Error (RDE) of 0.32 on the test set of the Development Dataset [2 ,3].