Binaural Audio-Visual Instance Segmentation

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of monaural audio-visual segmentation in distinguishing visually similar instances within the same category by introducing a novel task, Binaural Audio-Visual Instance Segmentation (BiAVIS), along with its corresponding benchmark, BiAVIS-Bench. Methodologically, this work incorporates physically grounded binaural spatial cues, extracting spatial priors through a sound source localization network. Furthermore, a query-level audio-visual fusion strategy coupled with an instance segmentation decoder is designed to effectively resolve intra-class ambiguity. Experimental results demonstrate that the proposed approach achieves substantial improvements on BiAVIS-Bench, increasing mAP by 17.22% and FSLA by 7.71%, significantly outperforming existing baselines.
📝 Abstract
Audio-visual segmentation (AVS) aims to segment sounding objects at the pixel level by integrating auditory and visual cues. However, existing methods are predominantly developed under the monaural setting and primarily rely on cross-modal semantic correspondence, which limits their ability to distinguish visually similar instances of the same semantic class. In contrast, humans naturally exploit binaural hearing, where interaural differences and direction-dependent acoustic filtering introduced by the head and pinnae provide physically grounded spatial cues for accurate sound source localization. Motivated by this observation, we introduce binaural audio-visual instance segmentation (BiAVIS), a new task that leverages synchronized binaural audio and video frames to segment sounding instances. To advance research on this task, we establish two benchmarks by manually annotating an existing binaural audio-visual dataset and collecting a new real-world dataset, BiAVIS-Bench, in more challenging and diverse scenarios. We further propose a BiAVIS model, which leverages an audio-only sound source localization network to learn spatial and semantic priors for sounding instances from binaural audio. A query-level audio-visual fusion strategy is subsequently introduced to inject these informative priors into the instance segmentation decoder. Extensive experiments conducted on the two proposed benchmarks demonstrate the superior performance of the BiAVIS model over previous monaural AVS methods, especially in resolving instance-level intra-class ambiguity. On the more challenging BiAVIS-Bench, the proposed BiAVIS model outperforms the best-performing monaural baselines by 17.22\% in mAP and 7.71\% in FSLA, respectively.
Problem

Research questions and friction points this paper is trying to address.

Binaural Audio-Visual Segmentation
Instance Segmentation
Sound Source Localization
Intra-class Ambiguity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Binaural Audio-Visual Instance Segmentation
Query-level Audio-Visual Fusion
Sound Source Localization
Spatial Priors
Instance-level Ambiguity
💼 Related Jobs
No related jobs found.
S
Saijun Wang
College of Electronic and Information Engineering, Tongji University, Shanghai 201804, China
G
Guanfeng Tang
College of Electronic and Information Engineering, Tongji University, Shanghai 201804, China
H
Hongbo Zhao
College of Electronic and Information Engineering, Tongji University, Shanghai 201804, China
Z
Zhicheng Lei
College of Electronic and Information Engineering, Tongji University, Shanghai 201804, China
Y
Yutong Zhang
College of Electronic and Information Engineering, Tongji University, Shanghai 201804, China
Wei Ye
Wei Ye
Tongji University
data miningmachine learningrepresentation learningdeep learning
R
Rui Fan
College of Electronic and Information Engineering, Shanghai Institute of Intelligent Science and Technology, State Key Laboratory of Autonomous Intelligent Unmanned Systems, Tongji University, Shanghai 201804, China