🤖 AI Summary
This study addresses three key challenges in world models: the absence of sound, real-time generation, and interactive control with spatial alignment. To this end, it proposes an interactive video-to-audio framework. Methodologically, we introduce a novel streaming causal autoregressive diffusion architecture that integrates chunk-indexed prompt scheduling, audio-centric caption guidance, and high-quality stereo supervision data to enable real-time spatial audio synthesis and dynamic manipulation. Experiments demonstrate that the proposed system achieves a real-time factor as low as 0.41, surpassing existing bidirectional models in both acoustic quality and spatial alignment on open-domain benchmarks. Ultimately, this work endows world models with a high-fidelity, interactive auditory dimension.
📝 Abstract
Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion. To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models. For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41. For interactive control, we incorporate an audio-centric captioning pipeline with chunk-indexed prompt scheduling, enabling dynamic manipulation of sound events during generation. For spatial alignment, we leverage high-quality stereo supervision curated from diverse stereo and ambisonic data. Extensive experiments demonstrate that while tailored for world models, WorldSonus generalizes effectively to open-domain video-to-audio benchmarks, matching or outperforming state-of-the-art bidirectional models in both acoustic quality and spatial alignment. Project page: https://noizai.github.io/WorldSonus/