geometry-conditioned view synthesis

Designs and implements algorithms and systems that synthesize geometrically consistent novel views and reconstruct 3D/4D scene structure from sparse multi-view or few-shot inputs using geometry-conditioned generation and rendering. This work builds geometry-constrained, multi-perspective embeddings and iterative refinement/reconstruction pipelines (including 3D-aware diffusion or iterative view-synthesis refinement) that recover occluded structure, augment limited-view coverage, and enforce cross-view and temporal consistency.

geometry-conditionedviewsynthesis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion

Jan 30, 2025
VG
Vitor Guizilini
🏛️ Toyota Research Institute | Toyota Technological Institute at Chicago

This work addresses novel view synthesis (NVS) from sparse multi-view images, proposing a zero-shot generative approach that avoids explicit 3D representations. Methodologically, it introduces: (1) a raymap-conditioned diffusion architecture that jointly models image and depth map generation; (2) learnable task embeddings for modality-aware conditional modulation; and (3) an efficient progressive fine-tuning paradigm where a compact model guides the adaptation of a large-scale diffusion model. The method integrates raymap spatial encoding, multi-task collaborative control, and large-scale multi-view data-driven learning. It achieves state-of-the-art performance across multiple NVS benchmarks and significantly improves accuracy and 3D consistency on downstream tasks—including multi-view stereo matching and video depth estimation—demonstrating strong generalization and geometric coherence without 3D supervision.

3D Scene ConsistencyDepth ImagingPerspective Generation

Geometry-Aware Diffusion Models for Multiview Scene Inpainting

Feb 18, 2025
AS
Ahmad Salimi
🏛️ York University | Samsung AI Centre Toronto | Google DeepMind | Vector Institute for AI

This work addresses 3D scene completion in occluded regions of multi-view images. We propose a geometry-aware conditional diffusion model that achieves high-fidelity, cross-view geometrically consistent image inpainting. Unlike methods relying on explicit/implicit radiance fields or requiring numerous input views, our approach jointly models geometric and appearance cues directly in the learned latent space. It incorporates a learnable implicit view-consistency constraint and geometry-guided cross-view feature alignment to circumvent blurring artifacts inherent in conventional fusion strategies. The method operates effectively under a few-view setting, substantially reducing data requirements. Evaluated on the SPIn-NeRF and NeRFiller benchmarks, it achieves state-of-the-art performance—particularly excelling in geometric consistency and fine-grained detail preservation.

3D scene inpainting with geometric consistencyAvoiding blurry images in multiview inpaintingFew-view inpainting with limited image sets

Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation

Jun 13, 2025
MK
Minseop Kwak
🏛️ NAVER AI Lab | KAIST AI | SNU AIIS

This paper addresses the alignment challenge in joint novel-view image and geometry generation under sparse reference images and coarse geometric priors. We propose the first diffusion-based image-geometry co-generation framework, formulated as cross-modal collaborative inpainting: an off-the-shelf pre-trained geometry predictor provides initial depth/normal maps; a novel cross-modal attention distillation mechanism enforces structural alignment between image and geometry diffusion branches; a proximity-aware mesh-conditioning strategy fuses depth and normal cues to suppress geometric noise; and geometry-guided warped inpainting enhances rendering consistency. Our method achieves high-fidelity extrapolative novel-view synthesis and synchronized geometry reconstruction on unseen scenes. It sets new state-of-the-art interpolation quality, outputs color-aligned point clouds with consistent geometry, and supports end-to-end 3D scene completion.

Aligned novel view image and geometry synthesisCross-modal attention for accurate alignmentHigh-fidelity extrapolative view synthesis

Controlling Space and Time with Diffusion Models

Jul 10, 2024
DW
Daniel Watson
🏛️ Google DeepMind

This work addresses the problem of 4D novel view synthesis (NVS) under arbitrary camera trajectories and timestamps, conditioned on single or multiple input images in natural scenes. To this end, we propose 4DiM—the first general-purpose 4D generative framework based on a cascaded diffusion architecture—breaking away from restrictive object-centric assumptions to support complex, open-world scenes. Methodologically, we introduce: (i) a structured-light motion reconstruction calibration pipeline enabling metric-scale pose control; (ii) a hybrid co-training paradigm integrating pose, time, and video modalities; and (iii) a conditional sampling mechanism for fine-grained spatiotemporal control. Experiments demonstrate that 4DiM significantly outperforms existing 3D NVS methods in both image fidelity and pose alignment accuracy. Moreover, it unifies diverse tasks—including single-image 3D generation, two-frame video interpolation/extrapolation, and pose-driven video-to-video translation—within a single framework. (149 words)

Enables 4D novel view synthesis with arbitrary camera trajectories and timestampsImproves generalization using mixed 3D, 4D, and video data trainingProvides intuitive metric-scale camera pose control for dynamic scenes

Latest Papers

What's happening recently
View more

This work addresses the challenge of balancing geometric consistency and camera controllability in large-baseline novel view synthesis from a single image. The authors propose a diffusion-based generative framework that integrates implicit geometric priors with sparse explicit metric depth cues. By leveraging a feedforward geometry-aware network to guide the generation process, the method achieves scale-consistent and structurally plausible view synthesis without requiring full 3D reconstruction. Experimental results demonstrate that the approach significantly outperforms existing methods under large viewpoint changes, exhibiting superior generalization capability and generation quality.

geometry consistencyimplicit geometrylarge viewpoint changes

This work addresses the challenge of simultaneously achieving high-fidelity 3D reconstruction and structurally plausible generation under sparse-view conditions. The authors propose a unified framework that, for the first time, cohesively integrates feed-forward reconstruction and diffusion-based generation across coordinate space, 3D representation, and training objectives. Key innovations include a shared canonical space alignment, decoupled collaborative learning, and an implicit geometric anchor-guided latent space enhancement mechanism. These components jointly enable efficient structural completion and consistent multi-view synthesis. The method significantly outperforms existing approaches under sparse observational settings, setting new state-of-the-art results in both reconstruction fidelity and robustness.

generative plausibilitymulti-view consistencyreconstruction fidelity

This work addresses the fundamental conflict in generative novel view synthesis—namely, the sparsity and inaccuracy of geometric priors coupled with the lack of geometric correspondence in appearance priors—by introducing a structured denoising dynamics mechanism. This approach achieves temporal decoupling and synergistic optimization of geometry and appearance during the diffusion process: early stages leverage geometric priors to establish a coarse structure, while later stages switch to appearance priors to correct geometric inaccuracies and refine fine details. Through a point-cloud-guided geometry-appearance fusion strategy, the method effectively disentangles these two components in both static and dynamic scenes, significantly outperforming existing approaches—particularly under severe point cloud sparsity or distortion—and enabling robust, high-quality novel view synthesis.

appearance priorsdiffusion processgeometric priors

Existing sparse-view 3D reconstruction methods suffer from limitations in geometric consistency and scene scalability, particularly struggling with large-scale or diverse scenes under arbitrary numbers of unordered inputs. To address these challenges, this work proposes AnyRecon, a novel framework that integrates persistent global 3D geometric memory with a geometry-aware conditioning mechanism to tightly couple generative and reconstructive processes. The approach introduces a pre-frame view cache to preserve inter-frame correspondences and leverages a video diffusion model enhanced with four-step distillation, sparse attention, and geometry-driven view retrieval. This design enables robust, high-quality, and scalable 3D reconstruction even under irregular inputs, large viewpoint gaps, and extended camera trajectories.

arbitrary-view reconstructiongeometric consistencylarge-scale 3D scenes

Hot Scholars

XZ

Xiatian Zhu

University of Surrey
Machine LearningComputer Vision
YZ

Yanyong Zhang

University of Science and Technology of China ; Rutgers University (Adjunct Visiting Professor)
SensingCyber-Physical SystemsMulti-Modal PerceptionEfficient AI Systems
TB

Thabo Beeler

Google
Digital Humans3D ReconstructionComputer GraphicsComputer Vision
MH

Marc Habermann

Senior Researcher, Max Planck Institute for Informatics
Computer VisionComputer GraphicsMachine LearningHuman Performance Capture
YL

Yangguang Li

CUHK
GenAIComputer GraphicsComputer Vision