text-driven scene synthesis

Designs and builds systems that convert natural-language prompts into structured scene layouts and 3D object arrangements, producing spatially consistent placements, poses, and inter-object relationships. Implements controllable generation and constraint handling to enforce specified semantics and spatial relationships while minimizing geometric conflicts and other layout violations.

text-drivenscenesynthesis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Layout-your-3D: Controllable and Precise 3D Generation with 2D Blueprint

Oct 20, 2024
JZ
Junwei Zhou
🏛️ Huazhong University of Science and Technology | NVIDIA | UC Merced | Yonsei University

Existing text-to-3D methods struggle to simultaneously ensure physically plausible object interactions and computational efficiency. This paper introduces Layout3D, a controllable and compositional 3D generation framework that leverages user-provided 2D layouts—optionally generated from text—as strong geometric priors. Its key contributions are: (1) the first end-to-end differentiable 3D generation paradigm guided by 2D layouts; (2) a collision-aware global layout optimization coupled with instance-level refinement, jointly ensuring structural physical plausibility and high-fidelity appearance; and (3) an integrated pipeline combining efficient reconstruction initialization, constraint-based optimization, instantiation rendering, and fine-tuning. Experiments demonstrate substantial improvements in geometric合理性 and visual fidelity of generated assets, with per-prompt inference time reduced by multiple orders of magnitude. Moreover, Layout3D natively supports downstream tasks such as 3D editing and object insertion.

Enables controllable 3D generation from text promptsImproves plausibility of object interactions in 3D assetsReduces optimization time for 3D generation processes

This work addresses the challenge that large language models face in comprehending spatial relationships in natural language and generating geometrically consistent layouts. The authors propose SG-Layout, a novel framework that explicitly incorporates structured scene graphs into large language models for the first time. By employing a graph encoder and a projector to align graph and language features, and leveraging LoRA for efficient instruction tuning while keeping the backbone network frozen, SG-Layout significantly enhances spatial reasoning accuracy and geometric consistency. The method demonstrates strong performance across diverse tasks—including image layout generation, indoor scene synthesis, and robotic object rearrangement—particularly excelling in scenarios characterized by dense relational structures and complex compositional arrangements.

geometric relationshipslarge language modelslayout generation

RoomCraft: Controllable and Complete 3D Indoor Scene Generation

Jun 27, 2025
MZ
Mengqi Zhou
🏛️ Chinese Academy of Sciences

This work addresses key challenges in 3D indoor scene generation from multimodal inputs (images, sketches, or text), including geometric inconsistency, frequent spatial conflicts, and incomplete layout synthesis. We propose a multi-stage generative framework that jointly incorporates semantic structure and spatial constraints. Our approach introduces a unified constraint representation and a conflict-aware localization strategy, integrates heuristic depth-first search (HDFS) to optimize object placement order, and designs a constraint-driven layout refinement module. To enhance semantic alignment, we incorporate natural language understanding into the generation pipeline. Experimental results demonstrate that our method consistently outperforms state-of-the-art approaches across all input modalities: furniture collision rate is reduced by 32.7%, and layout completeness improves by 28.4%. The generated scenes exhibit significantly enhanced spatial plausibility, visual realism, and semantic coherence.

Balancing geometric consistency and spatial relationshipsGenerating realistic 3D indoor scenes from user inputsMinimizing furniture collisions in multi-constraint scenarios

This work addresses the limited spatial understanding and layout consistency of current large language models and vision-language models in fine-grained visual editing. To overcome this, the authors propose a structured reasoning framework that explicitly models scene graph relationships to enable controllable and interpretable spatial layout editing guided by natural language instructions. The approach integrates scene graph representations, structured relational reasoning, and language guidance within a contrastive training paradigm, surpassing the limitations of conventional end-to-end methods and chain-of-thought supervised fine-tuning or GRPO strategies. Evaluated on a newly introduced benchmark for text-guided layout editing, the method achieves a 15% improvement in average IoU, reduces center distance error by 25%, and outperforms zero-shot state-of-the-art large language models by 20% in mIoU.

layout consistencyscene graphspatial reasoning

InstructLayout: Instruction-Driven 2D and 3D Layout Synthesis with Semantic Graph Prior

Jul 10, 2024
CL
Chenguo Lin
🏛️ Wangxuan Institute of Computer Technology | Peking University | PICO AI group | ByteDance

Existing natural language–driven 2D/3D layout generation methods rely on implicit modeling of object joint distributions and relational structures, resulting in poor controllability and low fidelity. To address this, we propose a semantic graph prior mechanism that explicitly decouples appearance representations from spatial distributions, enabling an instruction-conditioned layout decoder. We further integrate large language models (LLMs) and multimodal foundation models to automatically construct a high-quality, instruction–layout paired benchmark—constituting the first publicly available dataset of its kind. Our framework supports zero-shot generalization across tasks and dimensions (2D ↔ 3D). Experiments demonstrate significant improvements over state-of-the-art methods across multiple layout synthesis benchmarks. Ablation studies confirm the critical roles of the semantic graph prior and the co-designed data construction pipeline. Overall, our approach achieves highly controllable, high-fidelity, and dimensionally unified layout generation.

Addressing limitations in object relation modelingEnhancing controllability in 2D and 3D layout synthesisIntegrating semantic graph prior for improved fidelity

Latest Papers

What's happening recently
View more

ConsistCompose: Unified Multimodal Layout Control for Image Composition

Nov 23, 2025
XS
Xuanke Shi
🏛️ SenseTime Research

Existing unified multimodal models excel at visual understanding but suffer from significant limitations in layout-controllable multi-instance image generation (LELG), particularly in achieving precise compositional control. This paper introduces a novel Language-Embedded Layout Generation (LELG) paradigm: it directly encodes layout coordinates into language prompts and integrates coordinate-aware classifier-free guidance with a multimodal large model architecture, enabling joint text-image interleaved input and spatially accurate generation within a single unified interface—without task-specific branches—thus unifying multimodal understanding and generation. To support this, we construct the large-scale ConsistCompose3M dataset. Experiments demonstrate substantial improvements in spatial localization accuracy on COCO-Position and MS-Bench, while maintaining high identity fidelity. Our method achieves state-of-the-art performance in both multi-instance image generation and multimodal understanding.

Enabling layout-controlled multi-instance image generation from multimodal inputsImproving spatial accuracy while preserving identity fidelity in generationTranslating linguistic layout cues into precise spatial control mechanisms

This work addresses the challenge of error accumulation in reference frame transformations during multi-hop relative spatial reasoning, which often leads to 3D layouts that are inconsistent both semantically and metrically. To mitigate this issue, the authors propose the R³L framework, which decomposes spatial relationships into invariant subspaces to disentangle relational chains, employs an imagine-and-refine loop to enhance self-consistency, and reparameterizes coordinates from global to local to simplify pose optimization. By integrating multimodal large language models, spatial relation reasoning, and iterative refinement, R³L generates 3D layouts that better adhere to physical constraints and semantic coherence across diverse scenes and instructions, significantly alleviating the adverse effects of reference frame inconsistency in multi-hop spatial reasoning.

3D layout generationmulti-hop reasoningreference-frame transformation

Current large language models often produce erroneous layouts and object collisions in 3D indoor scene generation due to inadequate spatial representations. To address this, this work proposes SpatialGrammar—a compilable, domain-specific language (DSL) tailored for 3D indoor scenes—that encodes spatial structure through a gravity-aligned top-down grid representation and deterministically compiles into collision-free 3D geometry. Leveraging this DSL, we develop SG-Agent, a closed-loop optimization system, and SG-Mini, a lightweight 104M-parameter model, which together enable efficient scene generation trained exclusively on synthetic data for the first time. Experiments demonstrate that SG-Agent substantially improves spatial fidelity and physical plausibility, while SG-Mini matches or exceeds the performance of significantly larger LLM baselines in single-pass generation.

3D indoor scene generationcollision avoidancephysical plausibility

Hot Scholars

AR

Alireza Rastegarpanah

Co-founder of Extreme Robotics Lab, University of Birmingham
Robotic DisassemblyAI-driven RoboticsRobotic RemanufacturingMedical Robotics
JZ

Jiahuan Zhang

Imperial College London
Remote SensingDeep LearningGNSS
MP

Mattia Piccinini

TUM Global Post-doc Researcher, Technical University of Munich
Autonomous VehiclesArtificial IntelligenceRoboticsTrajectory Planning
ZC

Zhaopeng Cui

Zhejiang University
Computer VisionRoboticsComputer Graphics