🤖 AI Summary
This study addresses the degradation of extraction performance caused by error propagation from detection in acoustic scene segmentation. To mitigate this issue, we propose a multi-channel joint detection-extraction framework. The core innovation lies in an entropy-gated inference update mechanism that dynamically optimizes the inference process to effectively suppress cascading errors. Furthermore, confusion matrix relabeling and an audio judge model are introduced to enhance overall system robustness. Experimental evaluations on the DCASE 2025 dataset validate the effectiveness of each proposed module, demonstrating significant improvements in the synergistic performance of detection and extraction. This work establishes a solid foundation for practical applications in complex acoustic environments.
📝 Abstract
Spatial semantic segmentation of sound scenes (S5) involves the detection and extraction of target sound events in audio files. However, in typical detection-extraction pipelines, errors in the detection module propagate to the extraction module, degrading overall performance. In this work, we develop a multichannel detection-extraction model and evaluate inference-time algorithms to improve the class detection and target sound extraction in a multi-stage setup. We evaluate two fitness functions: binary cross entropy and mixture consistency; two relabeling strategies: naive and conditional confusion matrix relabeling; and three source estimate evaluation models: a multi-label classifier, a fine-tuned single-source classifier, and an off-the-shelf audio judge. We thoroughly evaluate the impact of these design choices on the DCASE 2025 Task 4 dataset, and validate entropy-based gating of inference-time updates on the held-out evaluation set, laying the groundwork for inference-time updates of S5 systems in real-world scenarios.