Generating Digital Models Using Text-to-3D and Image-to-3D Prompts: Critical Case Study
This work addresses the inconsistent modeling quality of current text-to-3D and image-to-3D generative tools. We present the first systematic, critical evaluation of mainstream online generators—including diffusion-based, NeRF-based, and implicit surface reconstruction methods—assessing their real-world output performance. Using a multi-prompt diversity benchmark coupled with a hybrid human-and-automated evaluation framework, we identify pervasive geometric distortions in complex topologies and fine-grained structures, and uncover key bottlenecks at the intersection of prompt engineering, geometric fidelity, and semantic consistency. Our contributions include actionable prompt design principles and a standardized set of quality evaluation metrics. These provide an empirical benchmark for applications such as digital twins and advance the development of next-generation generative 3D modeling technologies.