SAID: Semantic Acoustic Imaging Detector for Sound Event Localization and Detection
Existing sound source localization models typically estimate only directional information while neglecting spatial regions and energy distributions, thereby limiting their capacity for semantic acoustic imaging. To address this limitation, this work proposes the SAID model, which predicts labeled acoustic maps of active sound sources directly from audio, simultaneously characterizing their spatial extent, energy distribution, and categorical information. Methodologically, we introduce Audio2Sph, a novel pre-trained encoder that enhances audio representations through unsupervised sound energy estimation. Furthermore, a simulation-to-real fine-tuning pipeline is established, integrating self-supervised learning with multi-task optimization. The proposed model achieves first place in the DCASE2026 Task 3 Track A evaluation, attaining a Macro mAP of 0.1080 and a Pearson correlation coefficient of 0.3962.