Score
Designs and implements algorithms and systems that synthesize geometrically consistent novel views and reconstruct 3D/4D scene structure from sparse multi-view or few-shot inputs using geometry-conditioned generation and rendering. This work builds geometry-constrained, multi-perspective embeddings and iterative refinement/reconstruction pipelines (including 3D-aware diffusion or iterative view-synthesis refinement) that recover occluded structure, augment limited-view coverage, and enforce cross-view and temporal consistency.
Motivated by the urgent demands of AR/VR and digital twin applications for fast, generalizable, and deployment-friendly 3D reconstruction and novel view synthesis, this paper presents a systematic survey of feedforward deep learning methods—covering dominant representations including point clouds, 3D Gaussian splatting, and neural radiance fields—and focuses on three key challenges: pose-free input, dynamic scene modeling, and 3D-aware content generation. We propose the first unified taxonomy tailored to the feedforward paradigm, revealing inherent trade-offs between inference efficiency and cross-scene generalization. By integrating self-supervised learning, differentiable rendering, and multimodal input strategies—and leveraging standardized evaluation protocols and large-scale benchmarks—we comprehensively assess accuracy, latency, and robustness. Our analysis provides principled guidance and empirically grounded technology selection criteria for industrial-grade 3D vision systems.
This work addresses novel view synthesis (NVS) from sparse multi-view images, proposing a zero-shot generative approach that avoids explicit 3D representations. Methodologically, it introduces: (1) a raymap-conditioned diffusion architecture that jointly models image and depth map generation; (2) learnable task embeddings for modality-aware conditional modulation; and (3) an efficient progressive fine-tuning paradigm where a compact model guides the adaptation of a large-scale diffusion model. The method integrates raymap spatial encoding, multi-task collaborative control, and large-scale multi-view data-driven learning. It achieves state-of-the-art performance across multiple NVS benchmarks and significantly improves accuracy and 3D consistency on downstream tasks—including multi-view stereo matching and video depth estimation—demonstrating strong generalization and geometric coherence without 3D supervision.
This work addresses 3D scene completion in occluded regions of multi-view images. We propose a geometry-aware conditional diffusion model that achieves high-fidelity, cross-view geometrically consistent image inpainting. Unlike methods relying on explicit/implicit radiance fields or requiring numerous input views, our approach jointly models geometric and appearance cues directly in the learned latent space. It incorporates a learnable implicit view-consistency constraint and geometry-guided cross-view feature alignment to circumvent blurring artifacts inherent in conventional fusion strategies. The method operates effectively under a few-view setting, substantially reducing data requirements. Evaluated on the SPIn-NeRF and NeRFiller benchmarks, it achieves state-of-the-art performance—particularly excelling in geometric consistency and fine-grained detail preservation.
This paper addresses the alignment challenge in joint novel-view image and geometry generation under sparse reference images and coarse geometric priors. We propose the first diffusion-based image-geometry co-generation framework, formulated as cross-modal collaborative inpainting: an off-the-shelf pre-trained geometry predictor provides initial depth/normal maps; a novel cross-modal attention distillation mechanism enforces structural alignment between image and geometry diffusion branches; a proximity-aware mesh-conditioning strategy fuses depth and normal cues to suppress geometric noise; and geometry-guided warped inpainting enhances rendering consistency. Our method achieves high-fidelity extrapolative novel-view synthesis and synchronized geometry reconstruction on unseen scenes. It sets new state-of-the-art interpolation quality, outputs color-aligned point clouds with consistent geometry, and supports end-to-end 3D scene completion.
This work addresses the problem of 4D novel view synthesis (NVS) under arbitrary camera trajectories and timestamps, conditioned on single or multiple input images in natural scenes. To this end, we propose 4DiM—the first general-purpose 4D generative framework based on a cascaded diffusion architecture—breaking away from restrictive object-centric assumptions to support complex, open-world scenes. Methodologically, we introduce: (i) a structured-light motion reconstruction calibration pipeline enabling metric-scale pose control; (ii) a hybrid co-training paradigm integrating pose, time, and video modalities; and (iii) a conditional sampling mechanism for fine-grained spatiotemporal control. Experiments demonstrate that 4DiM significantly outperforms existing 3D NVS methods in both image fidelity and pose alignment accuracy. Moreover, it unifies diverse tasks—including single-image 3D generation, two-frame video interpolation/extrapolation, and pose-driven video-to-video translation—within a single framework. (149 words)
This work addresses the challenge of balancing geometric consistency and camera controllability in large-baseline novel view synthesis from a single image. The authors propose a diffusion-based generative framework that integrates implicit geometric priors with sparse explicit metric depth cues. By leveraging a feedforward geometry-aware network to guide the generation process, the method achieves scale-consistent and structurally plausible view synthesis without requiring full 3D reconstruction. Experimental results demonstrate that the approach significantly outperforms existing methods under large viewpoint changes, exhibiting superior generalization capability and generation quality.
This work addresses the challenge of simultaneously achieving high-fidelity 3D reconstruction and structurally plausible generation under sparse-view conditions. The authors propose a unified framework that, for the first time, cohesively integrates feed-forward reconstruction and diffusion-based generation across coordinate space, 3D representation, and training objectives. Key innovations include a shared canonical space alignment, decoupled collaborative learning, and an implicit geometric anchor-guided latent space enhancement mechanism. These components jointly enable efficient structural completion and consistent multi-view synthesis. The method significantly outperforms existing approaches under sparse observational settings, setting new state-of-the-art results in both reconstruction fidelity and robustness.
This work addresses the fundamental conflict in generative novel view synthesis—namely, the sparsity and inaccuracy of geometric priors coupled with the lack of geometric correspondence in appearance priors—by introducing a structured denoising dynamics mechanism. This approach achieves temporal decoupling and synergistic optimization of geometry and appearance during the diffusion process: early stages leverage geometric priors to establish a coarse structure, while later stages switch to appearance priors to correct geometric inaccuracies and refine fine details. Through a point-cloud-guided geometry-appearance fusion strategy, the method effectively disentangles these two components in both static and dynamic scenes, significantly outperforming existing approaches—particularly under severe point cloud sparsity or distortion—and enabling robust, high-quality novel view synthesis.
Existing sparse-view 3D reconstruction methods suffer from limitations in geometric consistency and scene scalability, particularly struggling with large-scale or diverse scenes under arbitrary numbers of unordered inputs. To address these challenges, this work proposes AnyRecon, a novel framework that integrates persistent global 3D geometric memory with a geometry-aware conditioning mechanism to tightly couple generative and reconstructive processes. The approach introduces a pre-frame view cache to preserve inter-frame correspondences and leverages a video diffusion model enhanced with four-step distillation, sparse attention, and geometry-driven view retrieval. This design enables robust, high-quality, and scalable 3D reconstruction even under irregular inputs, large viewpoint gaps, and extended camera trajectories.