SAID: Semantic Acoustic Imaging Detector for Sound Event Localization and Detection

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing sound source localization models typically estimate only directional information while neglecting spatial regions and energy distributions, thereby limiting their capacity for semantic acoustic imaging. To address this limitation, this work proposes the SAID model, which predicts labeled acoustic maps of active sound sources directly from audio, simultaneously characterizing their spatial extent, energy distribution, and categorical information. Methodologically, we introduce Audio2Sph, a novel pre-trained encoder that enhances audio representations through unsupervised sound energy estimation. Furthermore, a simulation-to-real fine-tuning pipeline is established, integrating self-supervised learning with multi-task optimization. The proposed model achieves first place in the DCASE2026 Task 3 Track A evaluation, attaining a Macro mAP of 0.1080 and a Pearson correlation coefficient of 0.3962.
📝 Abstract
In daily life, people hear speech, footsteps, and music around them. We can often recognize these sounds and judge where they come from. Each sound source can be shown on a separate acoustic map, a rectangular image covering $360^{\circ}$ horizontally and $180^{\circ}$ vertically. The map shows the directions occupied by the source as a region and the sound energy within that region. A class label identifies the sound. Predicting these labeled acoustic maps from audio is called semantic acoustic imaging. Such maps could help robots perceive their surroundings and allow augmented reality displays to show sound regions and classes over the real world. Existing models can recognize sound classes and estimate a direction for each source. However, a direction alone does not describe the source region or its energy. Acoustic imaging must also distinguish sound sources in nearby directions, while the number of active sources and the regions they occupy can change over time. We therefore propose the Semantic Acoustic Imaging Detector (SAID), which predicts a separate labeled acoustic map for each active source from audio. First, we pretrain Audio2Sph, SAID's audio encoder, through sound energy estimation across directions without class labels. Then, we train the complete SAID model to predict source regions, energy, and classes together. We also develop a pipeline that generates simulated recordings for pretraining and supports fine-tuning on real recordings. On the official DCASE2026 Task 3 Track A evaluation set, our submitted system ranks first with 0.1080 macro-averaged mean average precision (Macro mAP) and 0.3962 Macro Pearson $r$. Demos and code are provided at https://github.com/IN03X/SAID.
Problem

Research questions and friction points this paper is trying to address.

Sound Event Localization and Detection
Semantic Acoustic Imaging
Acoustic Map
Source Region Estimation
Sound Energy Estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Acoustic Imaging
Sound Event Localization and Detection
Audio2Sph
Pretraining
Simulated Recordings
🔎 Similar Papers
No similar papers found.
R
Runbang Wang
Nanjing University, China; The Chinese University of Hong Kong, Hong Kong SAR, China; Shun Hing Institute of Advanced Engineering (SHIAE), Hong Kong SAR, China
Z
Zining Liang
The Chinese University of Hong Kong, Hong Kong SAR, China; Shun Hing Institute of Advanced Engineering (SHIAE), Hong Kong SAR, China
Yin Cao
Yin Cao
Associate Professor, Xi'an Jiaotong-Liverpool University
Machine LearningAudio Signal ProcessingAcousticsNoise Control
Qiuqiang Kong
Qiuqiang Kong
The Chinese University of Hong Kong
Audio ProcessingArtificial Intelligence