WorldSonus: Bringing Sound to Worlds

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses three key challenges in world models: the absence of sound, real-time generation, and interactive control with spatial alignment. To this end, it proposes an interactive video-to-audio framework. Methodologically, we introduce a novel streaming causal autoregressive diffusion architecture that integrates chunk-indexed prompt scheduling, audio-centric caption guidance, and high-quality stereo supervision data to enable real-time spatial audio synthesis and dynamic manipulation. Experiments demonstrate that the proposed system achieves a real-time factor as low as 0.41, surpassing existing bidirectional models in both acoustic quality and spatial alignment on open-domain benchmarks. Ultimately, this work endows world models with a high-fidelity, interactive auditory dimension.
📝 Abstract
Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion. To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models. For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41. For interactive control, we incorporate an audio-centric captioning pipeline with chunk-indexed prompt scheduling, enabling dynamic manipulation of sound events during generation. For spatial alignment, we leverage high-quality stereo supervision curated from diverse stereo and ambisonic data. Extensive experiments demonstrate that while tailored for world models, WorldSonus generalizes effectively to open-domain video-to-audio benchmarks, matching or outperforming state-of-the-art bidirectional models in both acoustic quality and spatial alignment. Project page: https://noizai.github.io/WorldSonus/
Problem

Research questions and friction points this paper is trying to address.

world models
video-to-audio
real-time audio generation
interactive control
spatial audio
Innovation

Methods, ideas, or system contributions that make the work stand out.

world models
video-to-audio
causal autoregressive diffusion
interactive control
spatial audio
🔎 Similar Papers
No similar papers found.