Audible World Models: Spatially Aware Sound Generation for 3D Worlds

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of spatially aware audio in generative 3D worlds, where sound sources fail to adapt dynamically to listener movement. To overcome this limitation, we propose a training-free, end-to-end framework for generating spatial audio from text. This method is the first to explicitly correlate semantic layers, geometric structures, and acoustic propagation. By constructing panoramic 3D proxies and employing semantic stratification, it achieves robust sound source anchoring and geometric acoustic rendering, producing persistent, viewpoint-adaptive spatial audio without model fine-tuning. Evaluations across 80 scenes using vision-language models and human subjective assessments demonstrate that the proposed approach significantly enhances spatial consistency. Furthermore, users consistently preferred our method in terms of audiovisual coherence, spatial plausibility, and motion dependency.
📝 Abstract
Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synthesized solely from text or rendered video. Although such audio can convey what should be heard, it lacks an explicit representation of where sound sources are located and how their perceived sound should vary with listener movement. We introduce Audible World Models, a training-free framework that incorporates sound into the generated world state. Starting from a text prompt, our system constructs a panoramic 3D proxy, separates it into semantic layers, identifies sound-producing foreground objects and ambient background regions, and synthesizes dry audio for each sound label. It then anchors these sources to reconstructed geometry and renders listener-dependent spatial audio using geometric acoustic propagation. By explicitly linking semantics, geometry, and sound propagation, the framework maintains persistent source locations while adapting the rendered audio to changes in listener viewpoint and motion. Experiments across 80 generated scenes demonstrate substantial gains in spatial consistency over text-, video-, and panorama-conditioned baselines, while preserving competitive semantic alignment. VLM-based assessments and human evaluations further indicate that our soundtracks are preferred for their audio-visual consistency, spatial plausibility, and motion-dependent behavior.
Problem

Research questions and friction points this paper is trying to address.

3D world generation
spatial audio
sound source localization
audio-visual consistency
listener-dependent rendering
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Models
Spatial Audio
3D Scene Generation
Geometric Acoustic Propagation
Training-free Framework
🔎 Similar Papers