Score
Designs and builds systems that convert natural-language prompts into structured scene layouts and 3D object arrangements, producing spatially consistent placements, poses, and inter-object relationships. Implements controllable generation and constraint handling to enforce specified semantics and spatial relationships while minimizing geometric conflicts and other layout violations.
Existing text-to-3D methods struggle to simultaneously ensure physically plausible object interactions and computational efficiency. This paper introduces Layout3D, a controllable and compositional 3D generation framework that leverages user-provided 2D layouts—optionally generated from text—as strong geometric priors. Its key contributions are: (1) the first end-to-end differentiable 3D generation paradigm guided by 2D layouts; (2) a collision-aware global layout optimization coupled with instance-level refinement, jointly ensuring structural physical plausibility and high-fidelity appearance; and (3) an integrated pipeline combining efficient reconstruction initialization, constraint-based optimization, instantiation rendering, and fine-tuning. Experiments demonstrate substantial improvements in geometric合理性 and visual fidelity of generated assets, with per-prompt inference time reduced by multiple orders of magnitude. Moreover, Layout3D natively supports downstream tasks such as 3D editing and object insertion.
This work addresses the challenge that large language models face in comprehending spatial relationships in natural language and generating geometrically consistent layouts. The authors propose SG-Layout, a novel framework that explicitly incorporates structured scene graphs into large language models for the first time. By employing a graph encoder and a projector to align graph and language features, and leveraging LoRA for efficient instruction tuning while keeping the backbone network frozen, SG-Layout significantly enhances spatial reasoning accuracy and geometric consistency. The method demonstrates strong performance across diverse tasks—including image layout generation, indoor scene synthesis, and robotic object rearrangement—particularly excelling in scenarios characterized by dense relational structures and complex compositional arrangements.
This work addresses key challenges in 3D indoor scene generation from multimodal inputs (images, sketches, or text), including geometric inconsistency, frequent spatial conflicts, and incomplete layout synthesis. We propose a multi-stage generative framework that jointly incorporates semantic structure and spatial constraints. Our approach introduces a unified constraint representation and a conflict-aware localization strategy, integrates heuristic depth-first search (HDFS) to optimize object placement order, and designs a constraint-driven layout refinement module. To enhance semantic alignment, we incorporate natural language understanding into the generation pipeline. Experimental results demonstrate that our method consistently outperforms state-of-the-art approaches across all input modalities: furniture collision rate is reduced by 32.7%, and layout completeness improves by 28.4%. The generated scenes exhibit significantly enhanced spatial plausibility, visual realism, and semantic coherence.
This work addresses the limited spatial understanding and layout consistency of current large language models and vision-language models in fine-grained visual editing. To overcome this, the authors propose a structured reasoning framework that explicitly models scene graph relationships to enable controllable and interpretable spatial layout editing guided by natural language instructions. The approach integrates scene graph representations, structured relational reasoning, and language guidance within a contrastive training paradigm, surpassing the limitations of conventional end-to-end methods and chain-of-thought supervised fine-tuning or GRPO strategies. Evaluated on a newly introduced benchmark for text-guided layout editing, the method achieves a 15% improvement in average IoU, reduces center distance error by 25%, and outperforms zero-shot state-of-the-art large language models by 20% in mIoU.
Existing natural language–driven 2D/3D layout generation methods rely on implicit modeling of object joint distributions and relational structures, resulting in poor controllability and low fidelity. To address this, we propose a semantic graph prior mechanism that explicitly decouples appearance representations from spatial distributions, enabling an instruction-conditioned layout decoder. We further integrate large language models (LLMs) and multimodal foundation models to automatically construct a high-quality, instruction–layout paired benchmark—constituting the first publicly available dataset of its kind. Our framework supports zero-shot generalization across tasks and dimensions (2D ↔ 3D). Experiments demonstrate significant improvements over state-of-the-art methods across multiple layout synthesis benchmarks. Ablation studies confirm the critical roles of the semantic graph prior and the co-designed data construction pipeline. Overall, our approach achieves highly controllable, high-fidelity, and dimensionally unified layout generation.
针对3D空间文本到图像生成中对象关系、遮挡等问题,提出SpatialGuard框架,通过布局引导和视觉对齐等方法提高生成的可控性和准确性。
Existing unified multimodal models excel at visual understanding but suffer from significant limitations in layout-controllable multi-instance image generation (LELG), particularly in achieving precise compositional control. This paper introduces a novel Language-Embedded Layout Generation (LELG) paradigm: it directly encodes layout coordinates into language prompts and integrates coordinate-aware classifier-free guidance with a multimodal large model architecture, enabling joint text-image interleaved input and spatially accurate generation within a single unified interface—without task-specific branches—thus unifying multimodal understanding and generation. To support this, we construct the large-scale ConsistCompose3M dataset. Experiments demonstrate substantial improvements in spatial localization accuracy on COCO-Position and MS-Bench, while maintaining high identity fidelity. Our method achieves state-of-the-art performance in both multi-instance image generation and multimodal understanding.
This work addresses the challenge of error accumulation in reference frame transformations during multi-hop relative spatial reasoning, which often leads to 3D layouts that are inconsistent both semantically and metrically. To mitigate this issue, the authors propose the R³L framework, which decomposes spatial relationships into invariant subspaces to disentangle relational chains, employs an imagine-and-refine loop to enhance self-consistency, and reparameterizes coordinates from global to local to simplify pose optimization. By integrating multimodal large language models, spatial relation reasoning, and iterative refinement, R³L generates 3D layouts that better adhere to physical constraints and semantic coherence across diverse scenes and instructions, significantly alleviating the adverse effects of reference frame inconsistency in multi-hop spatial reasoning.
Current large language models often produce erroneous layouts and object collisions in 3D indoor scene generation due to inadequate spatial representations. To address this, this work proposes SpatialGrammar—a compilable, domain-specific language (DSL) tailored for 3D indoor scenes—that encodes spatial structure through a gravity-aligned top-down grid representation and deterministically compiles into collision-free 3D geometry. Leveraging this DSL, we develop SG-Agent, a closed-loop optimization system, and SG-Mini, a lightweight 104M-parameter model, which together enable efficient scene generation trained exclusively on synthetic data for the first time. Experiments demonstrate that SG-Agent substantially improves spatial fidelity and physical plausibility, while SG-Mini matches or exceeds the performance of significantly larger LLM baselines in single-pass generation.
研究解决了从单张图片生成可控和可执行3D场景的问题,通过统一视觉-语言-几何框架Fysiverse-3D-Vision实现空间推理与几何重建相互增强。