🤖 AI Summary
This study addresses the challenge of lacking explicit correspondence between geometric reconstruction and diffusion models in sparse-view novel view synthesis. To this end, we propose a geometry-routed multi-view diffusion model. Specifically, our method employs VGGT-Ω to extract 3D points along with their associated confidence scores, and introduces a novel visual geometry router that injects latent representations into a pretrained video diffusion model. Through a confidence-aware mechanism, this router guides the denoising process while preserving both foreground and background surface evidence. Furthermore, we incorporate point-trajectory residual consistency to enhance multi-view stability. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance on both interpolation and extrapolation tasks across varying levels of view difficulty. The source code has been made publicly available.
📝 Abstract
We present VGGT-Diff, a geometry-routed multi-view diffusion model for sparse-view novel view synthesis. Existing novel view synthesis (NVS) methods face a fundamental trade-off: reconstruction-based approaches preserve observed geometry but struggle to synthesize unseen regions, while diffusion-based methods provide strong generative priors yet rely on implicit source-to-query correspondence. VGGT-Diff bridges these regimes by routing visual geometry latents from VGGT-{\Omega} into a pretrained video diffusion model. Each visual token is associated with a 3D point and confidence, then transformed into query-aligned latent conditions through a confidence-aware Visual Geometry Router (VGR) that preserves front and back surface evidence. These conditions guide joint target-view denoising, while Point-Track Residual Consistency (PTRC) regularizes predicted-clean residuals along reliable 3D tracks, improving multi-view stability. We further introduce robust geometry conditioning, combining training-time regularization with inference-time guidance for improved robustness. Experiments show competitive or state-of-the-art performance across interpolation and extrapolation under different viewpoint difficulties. Our code is available at https://github.com/chenkangjie1123/VGGT-Diff.