🤖 AI Summary
This study addresses the difficulty existing generative models face in producing high-resolution, stereoscopic, and temporally consistent 4K 360° immersive videos. We propose a zero-shot generation pipeline that extends video diffusion models to this domain. Specifically, we design a binocular-vision-inspired, epipolar-aware 360° image matching metric to precisely capture cross-view temporal and stereo-geometric inconsistencies. Leveraging this metric as a training signal in conjunction with Direct Preference Optimization (DPO), our approach effectively resolves the model alignment challenge under limited data availability. The proposed method achieves high-quality 4K stereoscopic 360° video generation, providing a scalable content production pathway for mixed reality experiences.
📝 Abstract
Immersive displays can enable rich and diverse virtual experiences. Manually authoring every possible experience to realize this potential, however, is prohibitively expensive, difficult to scale, and impractical. Generative AI models could remove this bottleneck, but today's models are built for conventional displays and cannot generate the high-resolution, stereoscopic $360^\circ$ content required for immersive viewing. Further, temporal and stereo inconsistencies that may be tolerable on conventional displays can become highly disruptive when viewed through an immersive headset.
Here, we address this gap with a zero-shot generative pipeline that extends existing video diffusion models into 4K stereoscopic $360^\circ$ videos. Inspired from binocular vision and depth perception, we develop an epipolar-aware $360^\circ$ image matching metric that captures the temporal and stereo geometric inconsistencies across views. We then use this metric as a preference signal for direct preference optimization with limited training data. Our work enables $360^\circ$ stereo video generation and provides a scalable path for bringing generative content to immersive displays, allowing diverse mixed reality experiences on demand.