Optimized View and Geometry Distillation from Multi-view Diffuser

πŸ“… 2023-12-11
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses view inconsistency and geometric over-smoothing in single-view-to-multi-view image generation. We propose a radiance field optimization framework incorporating a consistency prior and unbiased score distillation (USD). First, we formulate radiance field optimization as a rigid geometric consistency priorβ€”novel in enforcing structural coherence across views. Second, USD corrects gradient bias inherent in conventional radiance field optimization, enabling more accurate geometry and appearance learning. Third, we design a two-stage diffusion model specialization pipeline that jointly optimizes object-specific priors and cross-view fidelity. Crucially, our method requires no large-scale multi-view training data and supports arbitrary camera poses. Experiments demonstrate state-of-the-art performance in multi-view synthesis and high-fidelity geometry-texture reconstruction, significantly improving view consistency, fine-detail recovery, and pose flexibility over existing approaches.
πŸ“ Abstract
Generating multi-view images from a single input view using image-conditioned diffusion models is a recent advancement and has shown considerable potential. However, issues such as the lack of consistency in synthesized views and over-smoothing in extracted geometry persist. Previous methods integrate multi-view consistency modules or impose additional supervisory to enhance view consistency while compromising on the flexibility of camera positioning and limiting the versatility of view synthesis. In this study, we consider the radiance field optimized during geometry extraction as a more rigid consistency prior, compared to volume and ray aggregation used in previous works. We further identify and rectify a critical bias in the traditional radiance field optimization process through score distillation from a multi-view diffuser. We introduce an Unbiased Score Distillation (USD) that utilizes unconditioned noises from a 2D diffusion model, greatly refining the radiance field fidelity. We leverage the rendered views from the optimized radiance field as the basis and develop a two-step specialization process of a 2D diffusion model, which is adept at conducting object-specific denoising and generating high-quality multi-view images. Finally, we recover faithful geometry and texture directly from the refined multi-view images. Empirical evaluations demonstrate that our optimized geometry and view distillation technique generates comparable results to the state-of-the-art models trained on extensive datasets, all while maintaining freedom in camera positioning. Please see our project page at https://youjiazhang.github.io/USD/.
Problem

Research questions and friction points this paper is trying to address.

Enhance multi-view consistency in synthesized images
Reduce over-smoothing in extracted geometry
Maintain camera positioning flexibility in view synthesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unbiased Score Distillation enhances radiance field fidelity
Two-step specialization improves multi-view image generation
Optimized radiance field ensures geometry and texture accuracy
πŸ”Ž Similar Papers
No similar papers found.
Huazhong University of Science and Technology
Y
Youjia Zhang
Huazhong University of Science and Technology
Junqing Yu
Junqing Yu
Huazhong University of Science & Technology
Zikai Song
Zikai Song
Huazhong University of Science and Technology
deep learningmultimedia modelsvisual trackingsocial media analysis
W
Wei Yang
Huazhong University of Science and Technology