UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

πŸ“… 2026-08-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of synthesizing geometrically consistent, high-fidelity, and camera-controllable novel views from extremely sparse input viewpointsβ€”a task where existing methods often fail. The authors propose UniWorld-View, a unified framework that synergistically combines explicit 3D geometric guidance with a video diffusion model to enable controllable large-baseline view synthesis from monocular images or videos. Central to this approach is an occlusion-aware point cloud rendering module that provides precise geometric priors tightly integrated with the diffusion process, effectively resolving geometric inconsistency and limited viewpoint control under sparse inputs. Experiments demonstrate that UniWorld-View significantly outperforms state-of-the-art methods on both WorldScore and zero-shot novel view synthesis benchmarks, achieving superior visual fidelity and controllability while supporting downstream dynamic 3D Gaussian splatting reconstruction.
πŸ“ Abstract
The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent views with precise camera control remains challenging when input coverage is extremely limited. Reconstruction-based approaches such as NeRF and 3D Gaussian Splatting (3DGS) deteriorate severely under sparse inputs and fail to explicitly handle occlusions. Generative methods ease data requirements but still struggle with large-baseline view synthesis due to inaccurate or implicit geometric guidance. To overcome these limitations, we introduce UniWorld-View, a unified framework for controllable large-baseline novel view synthesis from monocular inputs. UniWorld-View integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation. The geometric guidance is obtained through an occlusion-aware point cloud rendering strategy that resolves visibility ambiguities and provides accurate priors for diffusion-based synthesis. By coupling this rendering strategy with powerful video diffusion backbones, UniWorld-View achieves high-fidelity novel view generation even under extreme camera motions and wide-baseline changes, and can further provide multi-view videos for downstream dynamic 3DGS reconstruction. Experiments on the WorldScore benchmark and zero-shot NVS benchmarks demonstrate the effectiveness of UniWorld-View in controllability, geometric consistency, and visual fidelity.
Problem

Research questions and friction points this paper is trying to address.

novel view synthesis
large-baseline
monocular input
geometric consistency
occlusion handling
Innovation

Methods, ideas, or system contributions that make the work stand out.

large-baseline view synthesis
video diffusion models
explicit 3D guidance
occlusion-aware rendering
monocular novel view synthesis
πŸ”Ž Similar Papers
No similar papers found.