ReSpace: Text-Driven 3D Scene Synthesis and Editing with Preference Alignment

📅 2025-06-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing 3D indoor scene synthesis and editing methods suffer from semantic simplification (e.g., one-hot encoding), reliance on mask diffusion, neglect of room boundaries, constraints to rectangular layouts, or weak spatial reasoning. This paper proposes a text-driven, end-to-end 3D scene modeling framework. It introduces the first explicit room-boundary-guided, compact structured scene tokenization; employs a two-stage training strategy—supervised fine-tuning followed by preference alignment—to enhance instruction-following capability; and integrates a zero-shot LLM-coordinated editing mechanism with voxelized fine-grained geometric evaluation. The method supports non-rectangular free-form layouts, rich semantic understanding (e.g., “a modern studio with light-wood furniture”), and precise spatial editing. Experiments demonstrate significant improvements over state-of-the-art methods in object insertion tasks, while maintaining leading performance in full-scene synthesis.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Diffusion Models for VisionNatural Language Processing: Sentence-level Semantics, Textual Inference, etc.

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Scene synthesis and editing has emerged as a promising direction in computer graphics. Current trained approaches for 3D indoor scenes either oversimplify object semantics through one-hot class encodings (e.g., 'chair' or 'table'), require masked diffusion for editing, ignore room boundaries, or rely on floor plan renderings that fail to capture complex layouts. In contrast, LLM-based methods enable richer semantics via natural language (e.g., 'modern studio with light wood furniture') but do not support editing, remain limited to rectangular layouts or rely on weak spatial reasoning from implicit world models. We introduce ReSpace, a generative framework for text-driven 3D indoor scene synthesis and editing using autoregressive language models. Our approach features a compact structured scene representation with explicit room boundaries that frames scene editing as a next-token prediction task. We leverage a dual-stage training approach combining supervised fine-tuning and preference alignment, enabling a specially trained language model for object addition that accounts for user instructions, spatial geometry, object semantics, and scene-level composition. For scene editing, we employ a zero-shot LLM to handle object removal and prompts for addition. We further introduce a novel voxelization-based evaluation that captures fine-grained geometry beyond 3D bounding boxes. Experimental results surpass state-of-the-art on object addition while maintaining competitive results on full scene synthesis.
Problem

Research questions and friction points this paper is trying to address.

Simplify 3D scene synthesis with rich text semantics
Enable editable 3D scenes with explicit boundaries
Improve spatial reasoning for complex indoor layouts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Autoregressive language models for 3D scene synthesis
Dual-stage training with preference alignment
Zero-shot LLM for scene editing tasks