score-distillation 3d reconstruction

Design and implement methods that reconstruct 3D geometry and appearance by distilling score-based diffusion priors (e.g., via score distillation / SDS) into an optimization or generative pipeline, using gradients or guidance from 2D or 3D diffusion models to supervise novel-view synthesis. Build pipelines that jointly optimize geometry and texture, integrate reference-view supervision, and enforce multi-view 3D consistency while using diffusion-model guidance to shape the reconstruction.

score-distillation3dreconstruction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation

Jun 13, 2025
MK
Minseop Kwak
🏛️ NAVER AI Lab | KAIST AI | SNU AIIS

This paper addresses the alignment challenge in joint novel-view image and geometry generation under sparse reference images and coarse geometric priors. We propose the first diffusion-based image-geometry co-generation framework, formulated as cross-modal collaborative inpainting: an off-the-shelf pre-trained geometry predictor provides initial depth/normal maps; a novel cross-modal attention distillation mechanism enforces structural alignment between image and geometry diffusion branches; a proximity-aware mesh-conditioning strategy fuses depth and normal cues to suppress geometric noise; and geometry-guided warped inpainting enhances rendering consistency. Our method achieves high-fidelity extrapolative novel-view synthesis and synchronized geometry reconstruction on unseen scenes. It sets new state-of-the-art interpolation quality, outputs color-aligned point clouds with consistent geometry, and supports end-to-end 3D scene completion.

Aligned novel view image and geometry synthesisCross-modal attention for accurate alignmentHigh-fidelity extrapolative view synthesis

Optimized View and Geometry Distillation from Multi-view Diffuser

Dec 11, 2023
YZ
Youjia Zhang
🏛️ Huazhong University of Science and Technology

This work addresses view inconsistency and geometric over-smoothing in single-view-to-multi-view image generation. We propose a radiance field optimization framework incorporating a consistency prior and unbiased score distillation (USD). First, we formulate radiance field optimization as a rigid geometric consistency prior—novel in enforcing structural coherence across views. Second, USD corrects gradient bias inherent in conventional radiance field optimization, enabling more accurate geometry and appearance learning. Third, we design a two-stage diffusion model specialization pipeline that jointly optimizes object-specific priors and cross-view fidelity. Crucially, our method requires no large-scale multi-view training data and supports arbitrary camera poses. Experiments demonstrate state-of-the-art performance in multi-view synthesis and high-fidelity geometry-texture reconstruction, significantly improving view consistency, fine-detail recovery, and pose flexibility over existing approaches.

Enhance multi-view consistency in synthesized imagesMaintain camera positioning flexibility in view synthesisReduce over-smoothing in extracted geometry

Decompositional Neural Scene Reconstruction with Generative Diffusion Prior

Mar 19, 2025
JN
Junfeng Ni
🏛️ Tsinghua University | Peking University

Sparse-view 3D scene decomposition and reconstruction suffer from poor recovery of object-level geometric completeness and fine-grained texture—especially in under-constrained and occluded regions. To address this, we propose DP-Recon, the first framework to embed Score Distillation Sampling (SDS) diffusion priors into object-centric neural radiance field (NeRF) reconstruction. Our method introduces a visibility-guided dynamic weighting scheme to jointly optimize reconstruction fidelity and generative prior consistency. Furthermore, we design a visibility-aware SDS loss modulation strategy, significantly enhancing reconstruction fidelity and editability under sparse views. Evaluated on Replica and ScanNet++, DP-Recon substantially surpasses state-of-the-art methods: using only 10 input views, it outperforms baseline approaches trained on 100 views. It supports text-driven geometric and appearance editing and outputs VFX-ready meshes with high-fidelity UV parameterization.

Address underconstrained and occluded regions in sparse viewsDecompose 3D scenes with complete shapes and texturesEnhance geometry and appearance recovery using diffusion priors

Enhanced 3D Generation by 2D Editing

Dec 08, 2024
HL
Haoran Li
🏛️ University of Science and Technology of China | Nanyang Technological University | The Hong Kong University of Science and Technology | Aarhus University

Existing Score Distillation Sampling (SDS)-based 3D generation methods rely on single-step 2D denoising, leading to over-smoothed textures, impoverished geometric details, and limited content diversity. To address these limitations, we propose GE3D—a novel framework that formulates 3D generation as a multi-step latent-space 2D editing process. GE3D introduces a dual-trajectory alignment mechanism that jointly optimizes a noise-preserved fidelity trajectory and a text-guided denoising trajectory, enabling deep coupling between 3D representations and 2D diffusion priors. Through latent-space trajectory alignment, multi-granularity information distillation, and iterative editing refinement, GE3D significantly improves texture fidelity, geometric detail, and multi-view consistency of generated 3D assets. Quantitative and qualitative evaluations demonstrate state-of-the-art performance in material realism and text-3D alignment. The code and demo are publicly available.

Inefficient distillation from 2D diffusion models to 3D.Lack of rich content and diversity in 3D generation.Over-saturation and over-smoothing in current SDS-based methods.

3D-free meets 3D priors: Novel View Synthesis from a Single Image with Pretrained Diffusion Guidance

Aug 12, 2024
TK
Taewon Kang
🏛️ University of Maryland at College Park | Dolby Laboratories

Existing 3D novel view synthesis (NVS) methods rely heavily on dense 3D annotations or multi-view inputs, limiting generalization; conversely, 3D-free approaches enable text-driven generation of complex scenes but lack precise camera control. This paper proposes a single-image-driven, camera-controllable NVS framework. For the first time, it incorporates a pre-trained NVS model as a weak geometric prior into a 3D-free diffusion architecture and explicitly encodes camera pose in the CLIP embedding space. Built upon Stable Diffusion, our method integrates cross-modal CLIP enhancement, knowledge distillation, and pose-conditioned embedding—requiring neither multi-view nor 3D supervision. Evaluated on diverse real-world scenes, it significantly outperforms state-of-the-art methods, enabling accurate, arbitrary-view synthesis with high fidelity, geometric consistency, and fine-grained detail. Both qualitative and quantitative evaluations demonstrate leading performance.

Generating camera-controlled novel views from single input imagesHandling complex scenes without extensive 3D training dataOvercoming limitations of both 3D-based and 3D-free view synthesis methods

Latest Papers

What's happening recently
View more

Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single Images

Aug 04, 2025
PW
Philipp Wulff
🏛️ Technical University of Munich

This work addresses the challenging problem of high-fidelity 3D volumetric reconstruction from a single image without ground-truth 3D annotations or multi-view supervision. We propose a “diffusion-depth distillation” framework that leverages a pre-trained 2D diffusion model and a monocular depth estimator to generate geometric priors; implicit geometric knowledge is then distilled into a lightweight, feed-forward reconstruction network. To our knowledge, this is the first approach to jointly exploit 2D diffusion models and depth priors for monocular 3D reconstruction—eliminating reliance on costly 3D ground truth or multi-view data. Evaluated on KITTI-360 and Waymo, our method achieves performance on par with or surpassing state-of-the-art multi-view supervised methods, while demonstrating superior robustness and generalization, particularly in dynamic scenes.

Improving dynamic scene reconstruction accuracyMonocular 3D reconstruction from single imagesReducing dependency on expensive 3D supervision

RealisticDreamer: Guidance Score Distillation for Few-shot Gaussian Splatting

Nov 14, 2025
RW
Ruocheng Wu
🏛️ University of Electronic Science and Technology of China | Nanyang Technological University

To address overfitting in 3D Gaussian Splatting (3DGS) under sparse input views—caused by insufficient intermediate-view supervision—this paper proposes the first score-distillation guidance framework leveraging pre-trained video diffusion models. Methodologically, it introduces multi-view consistency priors encoded in video diffusion models into 3DGS optimization for the first time, and designs a unified guidance mechanism jointly utilizing depth warping and semantic features to rectify noise prediction directions, thereby mitigating score-distillation bias induced by motion and camera trajectory ambiguities. Additionally, multi-view rendering supervision is incorporated to enhance geometric accuracy and pose alignment. Experiments demonstrate that our approach significantly outperforms existing methods across multiple sparse-view datasets, achieving more robust 3D reconstruction and high-fidelity real-time rendering.

Addresses overfitting in sparse-view 3D Gaussian Splatting optimizationCorrects generative noise predictions using depth and semantic guidanceExtracts multi-view consistency priors from pretrained video diffusion models

This work addresses the challenge of insufficient geometric consistency and reconstruction fidelity in novel view synthesis under sparse observations and without camera pose information. The authors propose ReNoV, a framework that, for the first time, systematically leverages the geometric and semantic correspondences embedded in the spatial attention of external visual representations. By designing a dedicated representation projection module, ReNoV injects these correspondences as conditioning signals into a diffusion-based generative process, enabling high-quality view synthesis without explicit pose supervision. Evaluated on standard benchmarks, ReNoV significantly outperforms existing diffusion-based methods, achieving notable improvements in reconstruction fidelity, image inpainting quality, and geometric consistency.

diffusion modelgeometric consistencyhigh-fidelity generation

Single-view score distillation in 3D generation suffers from high gradient variance and global shape inconsistency. This work proposes Multi-View Score Distillation Interpolation (MV-SDI), which aggregates gradients from antipodal view pairs to substantially reduce gradient estimation variance along the camera axis, while keeping the pre-trained 2D diffusion model frozen, requiring no multi-view data, and incurring no additional peak memory cost. Under a fixed UNet evaluation budget, MV-SDI with only K=2 views improves CLIP R-Precision to 83.8% and halves the number of optimization steps; with K=4 views, it reduces the step count by fourfold and achieves an R-Precision of 86.9%, significantly outperforming single-view baselines across all alignment metrics.

3D GenerationGradient EstimationMulti-View Consistency

DiGA3D: Coarse-to-Fine Diffusional Propagation of Geometry and Appearance for Versatile 3D Inpainting

Jul 01, 2025
JP
Jingyi Pan
🏛️ The Hong Kong University of Science and Technology | The Hong Kong University of Science and Technology (Guangzhou)

This work addresses two critical limitations in existing 3D inpainting methods: poor robustness under large viewpoint variations in single-view approaches, and appearance/geometry inconsistency in multi-view diffusion-based repair. We propose the first unified 3D inpainting framework supporting object removal, retexuring, and replacement. Methodologically, we introduce a multi-reference view adaptive selection strategy, an attention-based feature propagation (AFP) mechanism to jointly optimize geometry and texture across views, and a texture-geometry joint score distillation sampling (TG-SDS) loss that explicitly enforces consistency between reconstructed geometry and surface appearance within repaired regions. Experiments demonstrate substantial improvements in cross-view consistency and robustness to large-angle reconstruction, effectively mitigating geometric distortion and visual incoherence—particularly under significant structural changes—where prior methods fail.

Appearance inconsistency from independent 2D diffusion-based inpaintingGeometry inconsistency due to significant inpainting region changesLack of robustness in single reference inpainting for distant views

Hot Scholars

YQ

Yue Qi

Beihang University
RZ

Ruiyuan Zhang

Zhejiang University
MultiModal3D Part AssemblyMixture-of-Expert
YZ

Yinghao Zhang

University of Cincinnati
Supply Chain ManagementBehavioral Operations Management
LM

Lennart Maack

Hamburg University of Technology
Deep LearningComputer VisionMedical Image AnalysisSurgical Data Science