🤖 AI Summary
Existing video-to-audio generation models lack dynamic spatial correspondence between stereo sound and visual content, limiting their applicability in immersive AR/VR environments. This work proposes StereoBind, a novel framework that introduces the concept of dynamic spatial correspondence to achieve spatial synergy between audio-visual signals via motion trajectory binding. Methodologically, it incorporates visual motion binding, a spatial trajectory encoder, and residual trajectory RoPE to precisely coordinate both modalities. To support this research, a large-scale dataset, StereoWorld-29K, and an evaluation benchmark, StereoWorldBench, are constructed. Experimental results demonstrate that the proposed framework significantly improves stereo spatial alignment accuracy while maintaining state-of-the-art overall audio-visual quality.
📝 Abstract
Recent joint video-audio generation models have achieved strong semantic correspondence and temporal synchronization. However, applications such as AR/VR and interactive gaming further require stereo audio to provide an immersive sense, which remains largely overlooked. Effective stereo audio requires the perceived sound location to evolve consistently with the motion of its corresponding visual source. We refer to this property as Dynamic Spatial Correspondence and propose StereoBind, a framework that binds visual source motion to stereo sound generation. StereoBind uses motion tracks to coordinate visual motion and stereo audio through three complementary mechanisms. Visual Motion Binding establishes source-aware audiovisual correspondence, the Spatial Track Encoder captures absolute source positions, and Residual Track RoPE models relative motion. For supervision and evaluation, we construct StereoWorld-29K, a large-scale stereo audio-video dataset with paired motion tracks, and StereoWorldBench for measuring audiovisual spatial consistency. Experiments show that StereoBind substantially improves spatial alignment in stereo audio generation over existing models while preserving overall audiovisual quality.