Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion

📅 2025-01-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses novel view synthesis (NVS) from sparse multi-view images, proposing a zero-shot generative approach that avoids explicit 3D representations. Methodologically, it introduces: (1) a raymap-conditioned diffusion architecture that jointly models image and depth map generation; (2) learnable task embeddings for modality-aware conditional modulation; and (3) an efficient progressive fine-tuning paradigm where a compact model guides the adaptation of a large-scale diffusion model. The method integrates raymap spatial encoding, multi-task collaborative control, and large-scale multi-view data-driven learning. It achieves state-of-the-art performance across multiple NVS benchmarks and significantly improves accuracy and 3D consistency on downstream tasks—including multi-view stereo matching and video depth estimation—demonstrating strong generalization and geometric coherence without 3D supervision.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Multi-instance/Multi-view LearningNatural Language Processing: Code Generation / Program Synthesis from Natural Language

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSocial Networks and Social Media: Generative AI / large language models and their impact on social systems
📝 Abstract
Current methods for 3D scene reconstruction from sparse posed images employ intermediate 3D representations such as neural fields, voxel grids, or 3D Gaussians, to achieve multi-view consistent scene appearance and geometry. In this paper we introduce MVGD, a diffusion-based architecture capable of direct pixel-level generation of images and depth maps from novel viewpoints, given an arbitrary number of input views. Our method uses raymap conditioning to both augment visual features with spatial information from different viewpoints, as well as to guide the generation of images and depth maps from novel views. A key aspect of our approach is the multi-task generation of images and depth maps, using learnable task embeddings to guide the diffusion process towards specific modalities. We train this model on a collection of more than 60 million multi-view samples from publicly available datasets, and propose techniques to enable efficient and consistent learning in such diverse conditions. We also propose a novel strategy that enables the efficient training of larger models by incrementally fine-tuning smaller ones, with promising scaling behavior. Through extensive experiments, we report state-of-the-art results in multiple novel view synthesis benchmarks, as well as multi-view stereo and video depth estimation.
Problem

Research questions and friction points this paper is trying to address.

Perspective Generation
Depth Imaging
3D Scene Consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-View Geometry Diffusion
View Synthesis
Depth Estimation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.