🤖 AI Summary
This study addresses the limited immersion of conventional videos caused by restricted fields of view and the absence of spatial audio. To this end, we propose OmniDream, a framework that transforms monocular silent videos into immersive audiovisual experiences supporting free-viewpoint navigation. The framework introduces an object-centric audio representation that innovatively decouples sound source content from acoustic properties. By integrating a training-free architecture with physics-based sound field propagation simulation, it achieves spatially precise alignment and flexible rendering of audio within visual scenes. Experimental results demonstrate that the proposed method significantly outperforms existing baselines in audiovisual alignment, spatial consistency, and perceived immersion.
📝 Abstract
Most videos capture only a narrow field of view and provide no spatial audio, limiting the sense of immersion they can provide. Recent video generation models can expand perspective videos into panoramic ones, but do not provide the corresponding spatial soundscape. Without spatially consistent audio, these expanded visual worlds remain incomplete. This paper presents OmniDream, a training-free framework that transforms a silent monocular video into an immersive audiovisual experience, where viewers can freely look around while sounds remain spatially aligned with the visual scene. At the core of OmniDream is an object-centric audio representation that disentangles each sound source's intrinsic audio content from its scene-dependent acoustic effects, enabling independent audio generation, physics-based simulation of propagation effects, and flexible spatial audio rendering. Experiments show improved audio-visual alignment, spatial correctness, and perceptual immersiveness over baselines. Examples are available on https://huggingface.co/spaces/CuriousAlien000/spatial-audio-360-demo