🤖 AI Summary
This study addresses the challenges of view inconsistency and limited geometric generalization in 3D material generation by proposing to fine-tune a video diffusion Transformer within texture space. By projecting images into texture space and leveraging pretrained priors, the method enables text-guided, multi-view, and super-resolution material generation and reconstruction that seamlessly adapts to arbitrary geometries. Notably, the proposed approach supports over one hundred input views and produces ultra-high-resolution outputs up to 8K. Extensive evaluations demonstrate that this framework achieves state-of-the-art performance across multiple material generation and reconstruction benchmarks.
📝 Abstract
We present a method for generating high quality materials for 3D objects entirely in texture space. We finetune a video diffusion transformer for text-guided material generation, multi-view material generation, and material upscaling. Our key insight is to use the known projection from image space to texture space, enabling the diffusion process to generalize across arbitrary geometries and texture parameterizations. This approach also avoids the view consistency issues inherent in video and multi-view diffusion models. Because texture space is two dimensional, we can reuse the strong priors of pretrained video diffusion models. We apply our method to high quality material reconstruction from posed photos captured under unknown lighting, as well as to text- and image guided material generation. Our method can scale to high resolutions (8K), 100+ input views, and neural material representations. In quantitative and qualitative evaluations we show state-of-the-art results for material generation and reconstruction.