🤖 AI Summary
This study addresses the challenges of training data scarcity, inadequate source-aware modeling, and the absence of standardized evaluation protocols in 3D asset editing by proposing Alchemy3D, a unified framework. Methodologically, we construct the first million-scale (1.25M) 3D editing dataset alongside an open-world benchmark, GEdit3D-Bench. The proposed architecture is built upon generative flow models, supporting multimodal text-image conditioning, few-step inference, and multi-view part segmentation transfer. Experimental results demonstrate that Alchemy3D significantly outperforms existing methods across core metrics, including editing fidelity, source consistency, and visual quality.
📝 Abstract
Although recent 3D generative models produce increasingly realistic assets, controllable 3D asset editing remains challenging. Existing methods are limited by scarce training data, insufficient source-aware modeling, and a lack of practical evaluation protocols. To address these limitations, we present Alchemy3D, a unified framework for training and evaluating versatile 3D asset editors that covers data construction, model architecture, and benchmark evaluation. Specifically, we curate Alchemy3D-1M, a large-scale 3D editing dataset containing 1.25M assets and 1.38M editing pairs across seven editing types. On this data, we train a family of generative flow models for general-purpose 3D asset editing. The model family supports image- and text-conditioned editing, few-step inference, and transfer to multi-view 3D part segmentation. We further introduce GEdit3D-Bench, a large-scale, open-world benchmark with a multi-dimensional evaluation protocol. Across existing and newly introduced benchmarks, our method outperforms prior methods on most metrics of editing fidelity, source preservation, and visual quality.