🤖 AI Summary
This study addresses the limited geometric completion capability of geometry-based methods and the cross-view inconsistency of video models in sparse-view synthesis. We propose a novel framework that performs generation within the geometric latent space of 3D foundation models. Specifically, we inject the Wan2.2 VACE video appearance prior via a ControlNet adapter and design a global spatial memory mechanism that reprojects to anchor a shared scene representation, pioneering the integration of video priors with spatial memory in the geometric latent space. This approach simultaneously preserves structural integrity and ensures cross-view consistency. Experimental results demonstrate that our method improves PSNR by 2.23 dB on the DL3DV dataset and reduces ATE by 32.4% on Mip-NeRF360, significantly enhancing both visual quality and geometric accuracy.
📝 Abstract
Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-based methods preserve observed scene structure but often struggle to complete unseen regions, whereas video generative models offer rich appearance priors but accumulate inconsistencies during sequential view generation. We propose GeoVerse, a framework that synthesizes world-consistent novel views by performing generation within the geometric latent space of a pretrained 3D foundation model and injecting appearance priors from a video generative model. Specifically, GeoVerse extracts multilevel features from Wan2.2 VACE and injects them into the geometric latent diffusion model via a ControlNet-style adapter, incorporating video-learned appearance priors to enhance structural completion. To enforce cross-view coherence, a global spatial memory continuously aggregates observed and synthesized content, reprojecting target-aligned guidance to anchor subsequent predictions to a shared scene representation. Extensive experiments across diverse datasets demonstrate improved visual quality and geometric consistency, with a 2.23 dB higher PSNR on DL3DV and 32.4% lower ATE on Mip-NeRF360 compared to GLD.