VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of lacking explicit correspondence between geometric reconstruction and diffusion models in sparse-view novel view synthesis. To this end, we propose a geometry-routed multi-view diffusion model. Specifically, our method employs VGGT-Ω to extract 3D points along with their associated confidence scores, and introduces a novel visual geometry router that injects latent representations into a pretrained video diffusion model. Through a confidence-aware mechanism, this router guides the denoising process while preserving both foreground and background surface evidence. Furthermore, we incorporate point-trajectory residual consistency to enhance multi-view stability. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance on both interpolation and extrapolation tasks across varying levels of view difficulty. The source code has been made publicly available.
📝 Abstract
We present VGGT-Diff, a geometry-routed multi-view diffusion model for sparse-view novel view synthesis. Existing novel view synthesis (NVS) methods face a fundamental trade-off: reconstruction-based approaches preserve observed geometry but struggle to synthesize unseen regions, while diffusion-based methods provide strong generative priors yet rely on implicit source-to-query correspondence. VGGT-Diff bridges these regimes by routing visual geometry latents from VGGT-{\Omega} into a pretrained video diffusion model. Each visual token is associated with a 3D point and confidence, then transformed into query-aligned latent conditions through a confidence-aware Visual Geometry Router (VGR) that preserves front and back surface evidence. These conditions guide joint target-view denoising, while Point-Track Residual Consistency (PTRC) regularizes predicted-clean residuals along reliable 3D tracks, improving multi-view stability. We further introduce robust geometry conditioning, combining training-time regularization with inference-time guidance for improved robustness. Experiments show competitive or state-of-the-art performance across interpolation and extrapolation under different viewpoint difficulties. Our code is available at https://github.com/chenkangjie1123/VGGT-Diff.
Problem

Research questions and friction points this paper is trying to address.

Sparse-View Novel View Synthesis
Reconstruction-based methods
Diffusion-based methods
Geometry preservation
Generative priors
Innovation

Methods, ideas, or system contributions that make the work stand out.

Novel View Synthesis
Multi-view Diffusion Model
Visual Geometry Router
Point-Track Residual Consistency
Sparse-View Reconstruction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Kangjie Chen
Kangjie Chen
Nanyang Technological University
Trustworthy AIRed-teamingBackdoor AttacksLLM-based Agents
X
Xiangyu Li
XPeng Motors
Dongbin Zhang
Dongbin Zhang
Tsinghua University
C
Chaoda Zheng
XPeng Motors
S
Shijia Chen
XPeng Motors
J
Jinhao Deng
XPeng Motors
H
Hongbin Lin
The Chinese University of Hong Kong
C
Choo Sin Wai
Tsinghua University
M
Minqi Wang
The Chinese University of Hong Kong
M
Minghao Yang
Tsinghua University
D
Dake Zhong
Tsinghua University
G
Guorui Song
Tsinghua University
Y
Yu Zhang
XPeng Motors
X
Xianming Liu
XPeng Motors
B
Boyang Wang
XPeng Motors