Score
Design and implement methods that reconstruct 3D geometry and appearance by distilling score-based diffusion priors (e.g., via score distillation / SDS) into an optimization or generative pipeline, using gradients or guidance from 2D or 3D diffusion models to supervise novel-view synthesis. Build pipelines that jointly optimize geometry and texture, integrate reference-view supervision, and enforce multi-view 3D consistency while using diffusion-model guidance to shape the reconstruction.
This survey addresses key challenges in 3D vision—occlusion robustness, point cloud sparsity, density imbalance, and high-dimensional computational bottlenecks—across four core tasks: 3D generation, point cloud reconstruction, shape completion, and scene synthesis. Methodologically, it introduces the first unified taxonomy capturing paradigm evolution, integrating denoising diffusion probabilistic models (DDPMs), 3D conditional encoders, multi-view feature alignment, implicit neural representations (INRs), and multimodal (text/image) guidance. The work rigorously delineates current performance limits and standardizes evaluation benchmarks. Crucially, it identifies three viable technical pathways forward: efficient sampling strategies, lightweight backward processes, and large-scale 3D pretraining. These contributions provide both theoretical foundations and practical guidelines for advancing diffusion-based 3D modeling.
This paper addresses the alignment challenge in joint novel-view image and geometry generation under sparse reference images and coarse geometric priors. We propose the first diffusion-based image-geometry co-generation framework, formulated as cross-modal collaborative inpainting: an off-the-shelf pre-trained geometry predictor provides initial depth/normal maps; a novel cross-modal attention distillation mechanism enforces structural alignment between image and geometry diffusion branches; a proximity-aware mesh-conditioning strategy fuses depth and normal cues to suppress geometric noise; and geometry-guided warped inpainting enhances rendering consistency. Our method achieves high-fidelity extrapolative novel-view synthesis and synchronized geometry reconstruction on unseen scenes. It sets new state-of-the-art interpolation quality, outputs color-aligned point clouds with consistent geometry, and supports end-to-end 3D scene completion.
This work addresses view inconsistency and geometric over-smoothing in single-view-to-multi-view image generation. We propose a radiance field optimization framework incorporating a consistency prior and unbiased score distillation (USD). First, we formulate radiance field optimization as a rigid geometric consistency prior—novel in enforcing structural coherence across views. Second, USD corrects gradient bias inherent in conventional radiance field optimization, enabling more accurate geometry and appearance learning. Third, we design a two-stage diffusion model specialization pipeline that jointly optimizes object-specific priors and cross-view fidelity. Crucially, our method requires no large-scale multi-view training data and supports arbitrary camera poses. Experiments demonstrate state-of-the-art performance in multi-view synthesis and high-fidelity geometry-texture reconstruction, significantly improving view consistency, fine-detail recovery, and pose flexibility over existing approaches.
Sparse-view 3D scene decomposition and reconstruction suffer from poor recovery of object-level geometric completeness and fine-grained texture—especially in under-constrained and occluded regions. To address this, we propose DP-Recon, the first framework to embed Score Distillation Sampling (SDS) diffusion priors into object-centric neural radiance field (NeRF) reconstruction. Our method introduces a visibility-guided dynamic weighting scheme to jointly optimize reconstruction fidelity and generative prior consistency. Furthermore, we design a visibility-aware SDS loss modulation strategy, significantly enhancing reconstruction fidelity and editability under sparse views. Evaluated on Replica and ScanNet++, DP-Recon substantially surpasses state-of-the-art methods: using only 10 input views, it outperforms baseline approaches trained on 100 views. It supports text-driven geometric and appearance editing and outputs VFX-ready meshes with high-fidelity UV parameterization.
Existing Score Distillation Sampling (SDS)-based 3D generation methods rely on single-step 2D denoising, leading to over-smoothed textures, impoverished geometric details, and limited content diversity. To address these limitations, we propose GE3D—a novel framework that formulates 3D generation as a multi-step latent-space 2D editing process. GE3D introduces a dual-trajectory alignment mechanism that jointly optimizes a noise-preserved fidelity trajectory and a text-guided denoising trajectory, enabling deep coupling between 3D representations and 2D diffusion priors. Through latent-space trajectory alignment, multi-granularity information distillation, and iterative editing refinement, GE3D significantly improves texture fidelity, geometric detail, and multi-view consistency of generated 3D assets. Quantitative and qualitative evaluations demonstrate state-of-the-art performance in material realism and text-3D alignment. The code and demo are publicly available.
Existing 3D novel view synthesis (NVS) methods rely heavily on dense 3D annotations or multi-view inputs, limiting generalization; conversely, 3D-free approaches enable text-driven generation of complex scenes but lack precise camera control. This paper proposes a single-image-driven, camera-controllable NVS framework. For the first time, it incorporates a pre-trained NVS model as a weak geometric prior into a 3D-free diffusion architecture and explicitly encodes camera pose in the CLIP embedding space. Built upon Stable Diffusion, our method integrates cross-modal CLIP enhancement, knowledge distillation, and pose-conditioned embedding—requiring neither multi-view nor 3D supervision. Evaluated on diverse real-world scenes, it significantly outperforms state-of-the-art methods, enabling accurate, arbitrary-view synthesis with high fidelity, geometric consistency, and fine-grained detail. Both qualitative and quantitative evaluations demonstrate leading performance.
This work addresses the challenging problem of high-fidelity 3D volumetric reconstruction from a single image without ground-truth 3D annotations or multi-view supervision. We propose a “diffusion-depth distillation” framework that leverages a pre-trained 2D diffusion model and a monocular depth estimator to generate geometric priors; implicit geometric knowledge is then distilled into a lightweight, feed-forward reconstruction network. To our knowledge, this is the first approach to jointly exploit 2D diffusion models and depth priors for monocular 3D reconstruction—eliminating reliance on costly 3D ground truth or multi-view data. Evaluated on KITTI-360 and Waymo, our method achieves performance on par with or surpassing state-of-the-art multi-view supervised methods, while demonstrating superior robustness and generalization, particularly in dynamic scenes.
To address overfitting in 3D Gaussian Splatting (3DGS) under sparse input views—caused by insufficient intermediate-view supervision—this paper proposes the first score-distillation guidance framework leveraging pre-trained video diffusion models. Methodologically, it introduces multi-view consistency priors encoded in video diffusion models into 3DGS optimization for the first time, and designs a unified guidance mechanism jointly utilizing depth warping and semantic features to rectify noise prediction directions, thereby mitigating score-distillation bias induced by motion and camera trajectory ambiguities. Additionally, multi-view rendering supervision is incorporated to enhance geometric accuracy and pose alignment. Experiments demonstrate that our approach significantly outperforms existing methods across multiple sparse-view datasets, achieving more robust 3D reconstruction and high-fidelity real-time rendering.
This work addresses the challenge of insufficient geometric consistency and reconstruction fidelity in novel view synthesis under sparse observations and without camera pose information. The authors propose ReNoV, a framework that, for the first time, systematically leverages the geometric and semantic correspondences embedded in the spatial attention of external visual representations. By designing a dedicated representation projection module, ReNoV injects these correspondences as conditioning signals into a diffusion-based generative process, enabling high-quality view synthesis without explicit pose supervision. Evaluated on standard benchmarks, ReNoV significantly outperforms existing diffusion-based methods, achieving notable improvements in reconstruction fidelity, image inpainting quality, and geometric consistency.
Single-view score distillation in 3D generation suffers from high gradient variance and global shape inconsistency. This work proposes Multi-View Score Distillation Interpolation (MV-SDI), which aggregates gradients from antipodal view pairs to substantially reduce gradient estimation variance along the camera axis, while keeping the pre-trained 2D diffusion model frozen, requiring no multi-view data, and incurring no additional peak memory cost. Under a fixed UNet evaluation budget, MV-SDI with only K=2 views improves CLIP R-Precision to 83.8% and halves the number of optimization steps; with K=4 views, it reduces the step count by fourfold and achieves an R-Precision of 86.9%, significantly outperforming single-view baselines across all alignment metrics.
This work addresses two critical limitations in existing 3D inpainting methods: poor robustness under large viewpoint variations in single-view approaches, and appearance/geometry inconsistency in multi-view diffusion-based repair. We propose the first unified 3D inpainting framework supporting object removal, retexuring, and replacement. Methodologically, we introduce a multi-reference view adaptive selection strategy, an attention-based feature propagation (AFP) mechanism to jointly optimize geometry and texture across views, and a texture-geometry joint score distillation sampling (TG-SDS) loss that explicitly enforces consistency between reconstructed geometry and surface appearance within repaired regions. Experiments demonstrate substantial improvements in cross-view consistency and robustness to large-angle reconstruction, effectively mitigating geometric distortion and visual incoherence—particularly under significant structural changes—where prior methods fail.