Score
Designs and implements algorithms and pipelines that procedurally synthesize re-renderable 3D scenes and datasets, producing assetized scene descriptions (geometry, materials, textures, lighting, and camera parameters), semantic/instance labels, and ground-truth 3D layouts. These systems enable photorealistic rendering across viewpoints and lighting conditions and support asset substitution, re-rendering, and large-scale automated dataset generation.
Manual 3D modeling remains labor-intensive and time-consuming, failing to meet the rapidly growing demands of XR and metaverse applications. Method: This paper presents a systematic survey of state-of-the-art methods for static 3D object and scene generation, introducing— for the first time—a multidimensional classification and cross-comparative framework that jointly considers representation evolution (e.g., point clouds, meshes, NeRFs) and generative paradigms (e.g., supervised learning, diffusion models, 2D foundation model priors, procedural modeling). Contribution/Results: We identify key challenges—including geometric-semantic consistency and scalability—and establish a reusable, multi-axis evaluation framework. Our analysis clarifies technological trajectories and performance boundaries, providing both theoretical foundations and practical guidelines for industrial-grade 3D content generation, thereby advancing the paradigm shift in 3D content creation.
3D scene generation suffers from limited diversity, low visual fidelity, and poor view consistency—hindering its deployment in immersive media, robotics, autonomous driving, and embodied AI. This paper presents a systematic survey of four dominant paradigms: procedural, neural 3D, image-driven, and video-driven generation. We introduce the first unified taxonomy to clarify technical evolution across these approaches. Three emerging frontiers are identified: physics-aware modeling, interactive generation, and perception-generation co-design. Leveraging NeRF, 3D Gaussian Splatting, diffusion models, GANs, and multimodal representations, we conduct a rigorous cross-paradigm evaluation on standard benchmarks, quantifying trade-offs among fidelity, diversity, and view consistency. To foster reproducibility and community advancement, we publicly release an open-source tracking platform that continuously monitors state-of-the-art progress.
Current text-to-3D generation suffers from heavy manual intervention, low efficiency, inconsistent stylistic outputs, and a lack of standardized evaluation protocols. Method: This paper proposes a data–architecture–evaluation co-optimization paradigm. We systematically survey generative AI techniques for 3D scene synthesis; introduce the first multi-dimensional evaluation framework tailored for text-to-3D generation; and integrate cross-attention mechanisms with latent-space alignment, augmented by multi-granularity metrics to quantify cross-modal alignment fidelity and data influence. Contribution/Results: We identify the core bottleneck in text–3D alignment; empirically validate the critical roles of data quality and architectural design in scalability; and establish a comprehensive benchmark balancing realism, stylistic controllability, and generation efficiency. Our work provides both theoretical foundations and practical methodologies for efficient, controllable, and stylistically consistent 3D content generation.
Existing 3D generation methods struggle to meet production-grade requirements for real-time interactive applications, such as consistent topology, UV unwrapping, physically based rendering (PBR) materials, skeletal rigging, and physically plausible scene layout. To address this gap, this work proposes a two-dimensional taxonomy centered on asset production pipelines—structured by asset type and production stage—and systematically constructs a comprehensive generation framework encompassing geometry synthesis, topology optimization, UV parameterization, PBR appearance modeling, skeletal rigging, and physics-aware scene assembly. The authors further introduce a cross-dimensional evaluation protocol to rigorously assess the direct usability of generated assets in game engines and simulation platforms. Their analysis highlights critical challenges in data quality, controllable generation, and end-to-end assetization, underscoring the pivotal role of deployable 3D content as foundational infrastructure for embodied intelligence and interactive world models.
Existing 3D generation methods struggle to simultaneously achieve high visual fidelity, real-time performance, and mobile deployment. This work proposes the first single-image 3D generation framework that balances deployment efficiency and interactive speed, producing high-quality meshes with baked normals, colored textures, and controllable face counts within 30 seconds; its Flash variant delivers preview-quality results in just 14 seconds. The approach integrates coarse-to-fine VecSet-based geometry generation, multi-view texture synthesis, and 3D back-projection inpainting, while performing mesh simplification, cleanup, normal baking, and parallel UV unwrapping directly on the GPU. Combined with model distillation and pipeline parallelism, the system minimizes end-to-end latency. Experiments demonstrate that the generated assets match the visual quality of commercial solutions, with both automated metrics and blind human evaluations confirming the method’s efficiency and practicality.
Existing 3D generative models neglect scene-specific constraints, resulting in synthetic assets that fail to integrate naturally into real-world environments. To address this, we propose a scene-conditioned framework for 3D object style transfer and compositing. Our method jointly optimizes object texture and environment lighting via differentiable ray tracing, while leveraging image priors from pre-trained text-to-image diffusion models (e.g., Stable Diffusion) to ensure geometric–photometric consistency and semantic adaptability in an end-to-end manner. Crucially, this work establishes the first tight coupling between 3D stylization and 2D scene semantics, enabling dynamic, context-aware relighting and restyling of a single 3D object across diverse semantic settings (e.g., summer/winter, fantasy/futuristic). We validate our approach on varied indoor/outdoor scenes and arbitrary 3D objects, demonstrating substantial improvements in visual realism, physical plausibility, and artistic controllability of composited imagery.
Single-image 3D scene generation with multiple objects faces severe challenges including heavy occlusion and object coupling-induced geometric distortion. To address these, we propose a two-stage differentiable framework: first, leveraging off-the-shelf image-to-3D models to independently reconstruct per-object meshes; second, jointly optimizing global scene layout via differentiable rendering, incorporating a novel optimal transport-driven long-range appearance loss and a high-level semantic loss in a synergistic constraint mechanism—enabling unified modeling of object-level geometric independence and scene-level structural consistency. Our approach integrates differentiable rendering, optimal transport theory, and semantics-guided gradient optimization. Evaluated on multi-object benchmarks, our method significantly improves geometric detail fidelity, object separation, and global coherence, outperforming state-of-the-art single-image 3D generation methods both quantitatively and qualitatively.
This work addresses key limitations in text-to-3D material generation—namely, heavy reliance on large-scale 3D-text paired data, limited editability, and insufficient photorealistic rendering fidelity. We propose an end-to-end framework that operates without 3D-text paired supervision. Our core innovations are threefold: (1) adopting procedural material graphs—not conventional texture maps—as the underlying material representation; (2) designing a segment-wise controlled diffusion model integrated with differentiable rendering to jointly optimize material parameters under text guidance; and (3) enabling fine-grained semantic control via geometric segmentation, text-guided 2D diffusion priors, and material graph parameter initialization. Experiments demonstrate substantial improvements over prior methods in realism, resolution, and interactive editability. The framework supports real-time, high-fidelity material synthesis and flexible, intuitive parameter adjustments—marking a significant step toward controllable, photorealistic text-driven material generation.
This work addresses the challenge of acquiring high-quality training data for multi-view stereo tasks, which is often costly and complex. The authors propose SimpleProc, a minimally rule-based procedural method for synthesizing multi-view image pairs through automated pipelines involving NURBS surfaces, procedural geometric modeling, displacement mapping, and texture synthesis. Remarkably, models trained on only 8,000 images generated by SimpleProc outperform those trained on an equivalent amount of real-world data. When scaled to 352,000 synthetic images, the approach matches or even exceeds the performance of models trained on 692,000 carefully curated real images, demonstrating the substantial efficiency and efficacy advantages of rule-driven synthetic data generation for multi-view stereo reconstruction.
Existing indoor scene generation methods rely on static meshes and predefined asset libraries, struggling to produce interactive, editable, and physically plausible objects. This work proposes a code-centric generative paradigm that frames scene construction as the synthesis of executable world programs: natural language prompts are automatically compiled into structured layouts and Blender Python scripts enriched with articulated joint metadata, enabling localized editing and state traceability. The approach integrates room-level agents, a plan-design-evaluate loop, five distinct code generation strategies, and an execution-guided repair mechanism, ultimately exporting simulation-ready scenes in SDF format. The resulting assets exhibit cleaner geometry and more accurate joint semantics, significantly outperforming existing methods in downstream tasks such as robotic interaction.
This work proposes the first end-to-end method for generating high-quality, editable 3D assets from a single image of indoor furniture or decor, tailored to the demands of interior design and e-commerce applications. The system employs a modular architecture comprising four coordinated stages—geometry reconstruction, texture generation, material assignment, and part decomposition—to produce watertight meshes with physically based rendering (PBR) materials and semantic part labels. Key innovations include implicit signed distance field (SDF) modeling via a geometry VAE and DiT, multi-view back-projection coupled with 3D texture field completion, MatWeaver-driven material matching, and multi-part joint SDF decoding enabled by PartVAE and PartDiT. Experiments demonstrate state-of-the-art performance across dedicated metrics, yielding high-fidelity 3D assets with accurate materials, structural completeness, and semantic editability.
This work proposes a method for editable 3D scene reconstruction from a single image that operates without requiring specialized 2D/3D foundation models, differentiable rendering, or multi-view supervision. By introducing an agent framework grounded in general-purpose vision-language models, the inverse graphics task is decomposed into staged optimization of geometry, materials, composition, and lighting, directly generating executable Blender scripts. This approach achieves, for the first time, high-quality, renderable, relightable, and controllable 3D scene reconstruction using only off-the-shelf vision-language models. It substantially improves fidelity at pixel, perceptual, and semantic levels across diverse scenes and enables a range of downstream editing and rendering applications.
This work addresses critical limitations in existing large language model (LLM)-based 3D scene generation methods for agricultural applications, including insufficient domain-specific knowledge, lack of validation mechanisms, and inadequate modularity, which collectively constrain controllability and scalability. To overcome these challenges, we propose a modular multi-LLM pipeline that integrates agricultural domain knowledge, few-shot prompting, retrieval-augmented generation (RAG), and Unreal Engine APIs to automatically construct realistic agricultural simulation environments. The architecture enables intermediate validation, structured data handling, and flexible extensibility, substantially enhancing semantic accuracy and visual fidelity. User studies and expert evaluations demonstrate that the system significantly outperforms manual design in both modeling efficiency and output quality, effectively overcoming the bottlenecks of conventional monolithic models in domain adaptation and controllable generation.