🤖 AI Summary
This work addresses key limitations in text-to-3D material generation—namely, heavy reliance on large-scale 3D-text paired data, limited editability, and insufficient photorealistic rendering fidelity. We propose an end-to-end framework that operates without 3D-text paired supervision. Our core innovations are threefold: (1) adopting procedural material graphs—not conventional texture maps—as the underlying material representation; (2) designing a segment-wise controlled diffusion model integrated with differentiable rendering to jointly optimize material parameters under text guidance; and (3) enabling fine-grained semantic control via geometric segmentation, text-guided 2D diffusion priors, and material graph parameter initialization. Experiments demonstrate substantial improvements over prior methods in realism, resolution, and interactive editability. The framework supports real-time, high-fidelity material synthesis and flexible, intuitive parameter adjustments—marking a significant step toward controllable, photorealistic text-driven material generation.
📝 Abstract
This paper aims to generate materials for 3D meshes from text descriptions. Unlike existing methods that synthesize texture maps, we propose to generate segment-wise procedural material graphs as the appearance representation, which supports high-quality rendering and provides substantial flexibility in editing. Instead of relying on extensive paired data, i.e., 3D meshes with material graphs and corresponding text descriptions, to train a material graph generative model, we propose to leverage the pre-trained 2D diffusion model as a bridge to connect the text and material graphs. Specifically, our approach decomposes a shape into a set of segments and designs a segment-controlled diffusion model to synthesize 2D images that are aligned with mesh parts. Based on generated images, we initialize parameters of material graphs and fine-tune them through the differentiable rendering module to produce materials in accordance with the textual description. Extensive experiments demonstrate the superior performance of our framework in photorealism, resolution, and editability over existing methods. Project page: https://zju3dv.github.io/MaPa