Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the mutual dependency between audio separation and sound source localization in multi-source audio-visual scenarios by proposing a two-stage self-supervised framework. Leveraging a selective convergence mechanism from contrastive learning, the method first localizes the dominant sound source and then iteratively uncovers additional sources using the initial localization as a prior, thereby mimicking human auditory attention. To mitigate existing evaluation biases, the study innovatively incorporates pixel-level semantic segmentation masks to establish a more equitable spatial alignment benchmark. On two-source benchmarks, the proposed approach substantially outperforms all existing self-supervised methods and even surpasses certain weakly supervised techniques on key metrics, advancing the state of the art in label-free multi-source sound localization.
📝 Abstract
Localising multiple sound sources in visual scenes remains a fundamental challenge in multimodal perception due to an inherent circular dependency: separating mixed audio requires knowing source locations, while identifying sound-producing regions requires separated audio signals. In this paper, we focus on the dual-source setting and discover a selective convergence in self-supervised audio-visual learning: when presented with multiple sound sources, contrastive models naturally converge to the most salient audio-visual correspondence rather than attempting to represent all sources equally. This emergent phenomenon, analogous to human selective auditory attention, enables us to break the above circular dependency through a progressive two-stage framework: first, leveraging selective convergence to identify dominant sources, and then exploiting these learned priors to uncover remaining sources. Our self-supervised approach achieves the best performance among self-supervised methods on dual-source benchmarks without requiring any manual annotations, and even surpasses some weakly-supervised approaches \red{on certain metrics. Furthermore, we identify a fundamental evaluation inconsistency in existing benchmarks: comparing continuous localisation heatmaps against bounding-box annotations creates systematic biases, particularly for non-axis-aligned objects where the bounding box includes substantial background regions. To address this, we introduce pixel-level segmentation masks to the existing benchmark, enabling spatially-aligned evaluation. Together, these results suggest that embracing rather than suppressing selectivity offers a scalable, annotation-free route to multi-source localisation.
Problem

Research questions and friction points this paper is trying to address.

audio-visual localisation
multi-source separation
circular dependency
self-supervised learning
selective attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

selective convergence
self-supervised learning
audio-visual localisation
multi-source separation
pixel-level evaluation