🤖 AI Summary
Existing interactive world models focus exclusively on visual rendering, neglecting synchronized acoustic dynamics. This work proposes the first real-time interactive audio-visual world model that enables the native co-evolution of visual scenes and spatial stereo sound under user interaction. Methodologically, it employs a bidirectional teacher pre-training framework with few-step streaming student distillation, and introduces an online trajectory distillation loss to ensure low-latency causal interaction. Additionally, a high-fidelity spatial audio-visual dataset and a spatial acoustic consistency evaluation benchmark are constructed. Experiments demonstrate that the model sustains drift-free joint generation at 24 FPS on a single GPU, achieving visual performance on par with state-of-the-art methods while significantly outperforming baselines in audio quality, thereby delivering a truly immersive experience.
📝 Abstract
World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive audio-visual world model where visual scenes and camera-grounded spatial stereo sound co-evolve natively under user interaction. We curate a high-fidelity spatial audio-visual dataset with true stereo acoustics and metric camera poses, upon which we pre-train a bidirectional teacher conditioned on 6-DoF camera trajectories and user actions. To enable low-latency causal interaction, we distill the teacher into a few-step streaming student via an online trajectory distillation loss, sustaining drift-free joint audio-visual rollouts at 24 FPS on a single GPU. Furthermore, we formalize spatial-acoustic consistency and introduce HelixBench to evaluate whether synthesized sound fields faithfully track dynamic viewpoint motion. Extensive experiments demonstrate that HelixWorld matches state-of-the-art silent world models in visual fidelity and responsiveness, while significantly surpassing existing baselines in camera-aligned spatial-acoustic immersion.