Supervising Sound Localization by In-the-wild Egomotion

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of effective supervision signals for binaural sound source localization. To overcome this limitation, we propose leveraging ego-motion as a form of weak supervision. Specifically, by employing multi-view geometry and camera motion estimation, visual ego-motion derived from in-the-wild videos is transformed into training signals for audio models, facilitating cross-modal learning when combined with binaural cues. The primary contribution of this work lies in being the first to exploit camera ego-motion from unconstrained videos to supervise sound source localization, alongside the construction of a real-world audiovisual dataset. Experimental results demonstrate that the proposed method achieves superior performance on the sound source localization task.
📝 Abstract
We present a method for learning binaural sound localization using egomotion as a supervisory signal. Over the course of a video, the cameras direction to a sound source will change as the camera moves. We train an audio model to predict sound directions that are consistent with visual estimates of camera motion, which we obtain using traditional methods from multi-view geometry. This provides a weak but plentiful form of supervision that we combine with traditional binaural cues. To evaluate this method, we propose a dataset of real-world audio-visual videos with egomotion. We show that our model can successfully learn from real-world data and that it performs well on sound localization tasks
Problem

Research questions and friction points this paper is trying to address.

Sound Localization
Binaural Audio
Egomotion
Audio-Visual Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

sound localization
egomotion
weak supervision
binaural cues
multi-view geometry
🔎 Similar Papers
No similar papers found.