procedural scene generation

Designs and implements algorithms and pipelines that procedurally synthesize re-renderable 3D scenes and datasets, producing assetized scene descriptions (geometry, materials, textures, lighting, and camera parameters), semantic/instance labels, and ground-truth 3D layouts. These systems enable photorealistic rendering across viewpoints and lighting conditions and support asset substitution, re-rendering, and large-scale automated dataset generation.

proceduralscenegeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.5
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

3D Scene Generation: A Survey

May 08, 2025
BW
Beichen Wen
🏛️ Nanyang Technological University

3D scene generation suffers from limited diversity, low visual fidelity, and poor view consistency—hindering its deployment in immersive media, robotics, autonomous driving, and embodied AI. This paper presents a systematic survey of four dominant paradigms: procedural, neural 3D, image-driven, and video-driven generation. We introduce the first unified taxonomy to clarify technical evolution across these approaches. Three emerging frontiers are identified: physics-aware modeling, interactive generation, and perception-generation co-design. Leveraging NeRF, 3D Gaussian Splatting, diffusion models, GANs, and multimodal representations, we conduct a rigorous cross-paradigm evaluation on standard benchmarks, quantifying trade-offs among fidelity, diversity, and view consistency. To foster reproducibility and community advancement, we publicly release an open-source tracking platform that continuously monitors state-of-the-art progress.

Addressing challenges in 3D representation and evaluationImproving fidelity and diversity with deep generative modelsSynthesizing structured 3D scenes for immersive applications

AI-powered Contextual 3D Environment Generation: A Systematic Review

Jun 05, 2025
MS
Miguel Silva
🏛️ University of Porto | INESC TEC

Current text-to-3D generation suffers from heavy manual intervention, low efficiency, inconsistent stylistic outputs, and a lack of standardized evaluation protocols. Method: This paper proposes a data–architecture–evaluation co-optimization paradigm. We systematically survey generative AI techniques for 3D scene synthesis; introduce the first multi-dimensional evaluation framework tailored for text-to-3D generation; and integrate cross-attention mechanisms with latent-space alignment, augmented by multi-granularity metrics to quantify cross-modal alignment fidelity and data influence. Contribution/Results: We identify the core bottleneck in text–3D alignment; empirically validate the critical roles of data quality and architectural design in scalability; and establish a comprehensive benchmark balancing realism, stylistic controllability, and generation efficiency. Our work provides both theoretical foundations and practical methodologies for efficient, controllable, and stylistically consistent 3D content generation.

Analyze challenges in scene authenticity and textual input influenceExplore training data impact and current model limitationsReview AI techniques for generating 3D environments efficiently

Must-Read Papers

Most classic and influential ideas
View more

Existing 3D generation methods struggle to meet production-grade requirements for real-time interactive applications, such as consistent topology, UV unwrapping, physically based rendering (PBR) materials, skeletal rigging, and physically plausible scene layout. To address this gap, this work proposes a two-dimensional taxonomy centered on asset production pipelines—structured by asset type and production stage—and systematically constructs a comprehensive generation framework encompassing geometry synthesis, topology optimization, UV parameterization, PBR appearance modeling, skeletal rigging, and physics-aware scene assembly. The authors further introduce a cross-dimensional evaluation protocol to rigorously assess the direct usability of generated assets in game engines and simulation platforms. Their analysis highlights critical challenges in data quality, controllable generation, and end-to-end assetization, underscoring the pivotal role of deployable 3D content as foundational infrastructure for embodied intelligence and interactive world models.

3D content generationasset usabilityengine-level constraints

Existing 3D generation methods struggle to simultaneously achieve high visual fidelity, real-time performance, and mobile deployment. This work proposes the first single-image 3D generation framework that balances deployment efficiency and interactive speed, producing high-quality meshes with baked normals, colored textures, and controllable face counts within 30 seconds; its Flash variant delivers preview-quality results in just 14 seconds. The approach integrates coarse-to-fine VecSet-based geometry generation, multi-view texture synthesis, and 3D back-projection inpainting, while performing mesh simplification, cleanup, normal baking, and parallel UV unwrapping directly on the GPU. Combined with model distillation and pipeline parallelism, the system minimizes end-to-end latency. Experiments demonstrate that the generated assets match the visual quality of commercial solutions, with both automated metrics and blind human evaluations confirming the method’s efficiency and practicality.

3D asset generationdeployabilityinteractive speed

Scene-Conditional 3D Object Stylization and Composition

Dec 19, 2023
JZ
Jinghao Zhou
🏛️ University of Oxford

Existing 3D generative models neglect scene-specific constraints, resulting in synthetic assets that fail to integrate naturally into real-world environments. To address this, we propose a scene-conditioned framework for 3D object style transfer and compositing. Our method jointly optimizes object texture and environment lighting via differentiable ray tracing, while leveraging image priors from pre-trained text-to-image diffusion models (e.g., Stable Diffusion) to ensure geometric–photometric consistency and semantic adaptability in an end-to-end manner. Crucially, this work establishes the first tight coupling between 3D stylization and 2D scene semantics, enabling dynamic, context-aware relighting and restyling of a single 3D object across diverse semantic settings (e.g., summer/winter, fantasy/futuristic). We validate our approach on varied indoor/outdoor scenes and arbitrary 3D objects, demonstrating substantial improvements in visual realism, physical plausibility, and artistic controllability of composited imagery.

Adapting object appearance to environmental changesEnhancing object-scene composition realismStylizing 3D objects to match 2D scenes

REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout Alignment

May 28, 2024
HH
Haonan Han
🏛️ Tsinghua University | The University of Hong Kong | Tencent Meeting | Harvard University

Single-image 3D scene generation with multiple objects faces severe challenges including heavy occlusion and object coupling-induced geometric distortion. To address these, we propose a two-stage differentiable framework: first, leveraging off-the-shelf image-to-3D models to independently reconstruct per-object meshes; second, jointly optimizing global scene layout via differentiable rendering, incorporating a novel optimal transport-driven long-range appearance loss and a high-level semantic loss in a synergistic constraint mechanism—enabling unified modeling of object-level geometric independence and scene-level structural consistency. Our approach integrates differentiable rendering, optimal transport theory, and semantics-guided gradient optimization. Evaluated on multi-object benchmarks, our method significantly improves geometric detail fidelity, object separation, and global coherence, outperforming state-of-the-art single-image 3D generation methods both quantitatively and qualitatively.

Addresses multi-object scene biases and occlusion complexitiesGenerates compositional 3D assets from single imagesOptimizes 3D layout alignment through differentiable rendering

MaPa: Text-driven Photorealistic Material Painting for 3D Shapes

Apr 26, 2024
SZ
Shangzhan Zhang
🏛️ Zhejiang University | Ant Group | Shenzhen University

This work addresses key limitations in text-to-3D material generation—namely, heavy reliance on large-scale 3D-text paired data, limited editability, and insufficient photorealistic rendering fidelity. We propose an end-to-end framework that operates without 3D-text paired supervision. Our core innovations are threefold: (1) adopting procedural material graphs—not conventional texture maps—as the underlying material representation; (2) designing a segment-wise controlled diffusion model integrated with differentiable rendering to jointly optimize material parameters under text guidance; and (3) enabling fine-grained semantic control via geometric segmentation, text-guided 2D diffusion priors, and material graph parameter initialization. Experiments demonstrate substantial improvements over prior methods in realism, resolution, and interactive editability. The framework supports real-time, high-fidelity material synthesis and flexible, intuitive parameter adjustments—marking a significant step toward controllable, photorealistic text-driven material generation.

Create segment-wise procedural material graphs for editingGenerate 3D mesh materials from text descriptionsLeverage 2D diffusion models without paired 3D training data

Latest Papers

What's happening recently
View more

This work addresses the challenge of acquiring high-quality training data for multi-view stereo tasks, which is often costly and complex. The authors propose SimpleProc, a minimally rule-based procedural method for synthesizing multi-view image pairs through automated pipelines involving NURBS surfaces, procedural geometric modeling, displacement mapping, and texture synthesis. Remarkably, models trained on only 8,000 images generated by SimpleProc outperform those trained on an equivalent amount of real-world data. When scaled to 352,000 synthetic images, the approach matches or even exceeds the performance of models trained on 692,000 carefully curated real images, demonstrating the substantial efficiency and efficacy advantages of rule-driven synthetic data generation for multi-view stereo reconstruction.

multi-view stereoNURBSprocedural generation

Existing indoor scene generation methods rely on static meshes and predefined asset libraries, struggling to produce interactive, editable, and physically plausible objects. This work proposes a code-centric generative paradigm that frames scene construction as the synthesis of executable world programs: natural language prompts are automatically compiled into structured layouts and Blender Python scripts enriched with articulated joint metadata, enabling localized editing and state traceability. The approach integrates room-level agents, a plan-design-evaluate loop, five distinct code generation strategies, and an execution-guided repair mechanism, ultimately exporting simulation-ready scenes in SDF format. The resulting assets exhibit cleaner geometry and more accurate joint semantics, significantly outperforming existing methods in downstream tasks such as robotic interaction.

articulated objectscontrollabilityexecutable programs

This work proposes the first end-to-end method for generating high-quality, editable 3D assets from a single image of indoor furniture or decor, tailored to the demands of interior design and e-commerce applications. The system employs a modular architecture comprising four coordinated stages—geometry reconstruction, texture generation, material assignment, and part decomposition—to produce watertight meshes with physically based rendering (PBR) materials and semantic part labels. Key innovations include implicit signed distance field (SDF) modeling via a geometry VAE and DiT, multi-view back-projection coupled with 3D texture field completion, MatWeaver-driven material matching, and multi-part joint SDF decoding enabled by PartVAE and PartDiT. Experiments demonstrate state-of-the-art performance across dedicated metrics, yielding high-fidelity 3D assets with accurate materials, structural completeness, and semantic editability.

3D asset generationfurniture modelingimage-to-3D

This work proposes a method for editable 3D scene reconstruction from a single image that operates without requiring specialized 2D/3D foundation models, differentiable rendering, or multi-view supervision. By introducing an agent framework grounded in general-purpose vision-language models, the inverse graphics task is decomposed into staged optimization of geometry, materials, composition, and lighting, directly generating executable Blender scripts. This approach achieves, for the first time, high-quality, renderable, relightable, and controllable 3D scene reconstruction using only off-the-shelf vision-language models. It substantially improves fidelity at pixel, perceptual, and semantic levels across diverse scenes and enables a range of downstream editing and rendering applications.

3D scene reconstructionBlenderexecutable representation

This work addresses critical limitations in existing large language model (LLM)-based 3D scene generation methods for agricultural applications, including insufficient domain-specific knowledge, lack of validation mechanisms, and inadequate modularity, which collectively constrain controllability and scalability. To overcome these challenges, we propose a modular multi-LLM pipeline that integrates agricultural domain knowledge, few-shot prompting, retrieval-augmented generation (RAG), and Unreal Engine APIs to automatically construct realistic agricultural simulation environments. The architecture enables intermediate validation, structured data handling, and flexible extensibility, substantially enhancing semantic accuracy and visual fidelity. User studies and expert evaluations demonstrate that the system significantly outperforms manual design in both modeling efficiency and output quality, effectively overcoming the bottlenecks of conventional monolithic models in domain adaptation and controllable generation.

3D scene generationagricultural simulationdomain-specific reasoning

Hot Scholars

YZ

Yuke Zhu

The University of Texas at Austin, NVIDIA Research
Robot LearningComputer VisionMachine LearningRobotics
RK

Ranjay Krishna

University of Washington, Allen Institute for AI
Computer VisionNatural Language ProcessingMachine LearningHuman Computer Interaction
SZ

Sergey Zakharov

Toyota Research Institute
Computer VisionMachine LearningAugmented Reality
JD

Jiafei Duan

Computer Science PhD Student, University of Washington
RoboticsRobot LearningEmbodied AIRobotic Manipulation
RH

Rose Hendrix

Research Engineer @ PRIOR, AI2
roboticsmachine learning