🤖 AI Summary
Current text-to-3D generation suffers from heavy manual intervention, low efficiency, inconsistent stylistic outputs, and a lack of standardized evaluation protocols.
Method: This paper proposes a data–architecture–evaluation co-optimization paradigm. We systematically survey generative AI techniques for 3D scene synthesis; introduce the first multi-dimensional evaluation framework tailored for text-to-3D generation; and integrate cross-attention mechanisms with latent-space alignment, augmented by multi-granularity metrics to quantify cross-modal alignment fidelity and data influence.
Contribution/Results: We identify the core bottleneck in text–3D alignment; empirically validate the critical roles of data quality and architectural design in scalability; and establish a comprehensive benchmark balancing realism, stylistic controllability, and generation efficiency. Our work provides both theoretical foundations and practical methodologies for efficient, controllable, and stylistically consistent 3D content generation.
📝 Abstract
The generation of high-quality 3D environments is crucial for industries such as gaming, virtual reality, and cinema, yet remains resource-intensive due to the reliance on manual processes. This study performs a systematic review of existing generative AI techniques for 3D scene generation, analyzing their characteristics, strengths, limitations, and potential for improvement. By examining state-of-the-art approaches, it presents key challenges such as scene authenticity and the influence of textual inputs. Special attention is given to how AI can blend different stylistic domains while maintaining coherence, the impact of training data on output quality, and the limitations of current models. In addition, this review surveys existing evaluation metrics for assessing realism and explores how industry professionals incorporate AI into their workflows. The findings of this study aim to provide a comprehensive understanding of the current landscape and serve as a foundation for future research on AI-driven 3D content generation. Key findings include that advanced generative architectures enable high-quality 3D content creation at a high computational cost, effective multi-modal integration techniques like cross-attention and latent space alignment facilitate text-to-3D tasks, and the quality and diversity of training data combined with comprehensive evaluation metrics are critical to achieving scalable, robust 3D scene generation.