🤖 AI Summary
This study addresses the challenges of foreground geometric ambiguity and insufficient spatial coherence in single-image novel view synthesis by proposing a novel framework based on decoupled 3D priors. Methodologically, 3D Gaussian Splatting is employed to separately model the foreground and background, followed by a coarse-to-fine geometric optimization and alignment procedure. Concurrently, a dual-stream masking mechanism is designed to guide video diffusion models in generating high-quality novel views. This work not only significantly enhances both visual quality and geometric accuracy, outperforming existing mainstream methods, but also inherently supports flexible scene editing capabilities owing to its decoupled representation.
📝 Abstract
Single-image novel view synthesis remains challenging because the underlying 3D geometry is highly ambiguous. Recent diffusion-based approaches produce plausible results, but they often struggle to preserve the geometric structure and spatial coherence of foreground objects. We present GenNVS, a framework for geometry-enhanced novel view synthesis via a disentangled 3D prior. Specifically, GenNVS models foreground objects and the background with 3D Gaussian Splatting and aligns them through a coarse-to-fine geometric optimization process to form a unified 3D scene. This scene conditions a video diffusion model through the proposed Dual-Stream Masking mechanism, which guides synthesis by jointly exploiting rendered validity masks and geometry-aware warping. Experimental results show that GenNVS performs favorably against recent methods in both visual quality and geometric accuracy, while naturally supporting flexible scene editing.