🤖 AI Summary
Existing 3D indoor scene synthesis and editing methods suffer from semantic simplification (e.g., one-hot encoding), reliance on mask diffusion, neglect of room boundaries, constraints to rectangular layouts, or weak spatial reasoning. This paper proposes a text-driven, end-to-end 3D scene modeling framework. It introduces the first explicit room-boundary-guided, compact structured scene tokenization; employs a two-stage training strategy—supervised fine-tuning followed by preference alignment—to enhance instruction-following capability; and integrates a zero-shot LLM-coordinated editing mechanism with voxelized fine-grained geometric evaluation. The method supports non-rectangular free-form layouts, rich semantic understanding (e.g., “a modern studio with light-wood furniture”), and precise spatial editing. Experiments demonstrate significant improvements over state-of-the-art methods in object insertion tasks, while maintaining leading performance in full-scene synthesis.
📝 Abstract
Scene synthesis and editing has emerged as a promising direction in computer graphics. Current trained approaches for 3D indoor scenes either oversimplify object semantics through one-hot class encodings (e.g., 'chair' or 'table'), require masked diffusion for editing, ignore room boundaries, or rely on floor plan renderings that fail to capture complex layouts. In contrast, LLM-based methods enable richer semantics via natural language (e.g., 'modern studio with light wood furniture') but do not support editing, remain limited to rectangular layouts or rely on weak spatial reasoning from implicit world models. We introduce ReSpace, a generative framework for text-driven 3D indoor scene synthesis and editing using autoregressive language models. Our approach features a compact structured scene representation with explicit room boundaries that frames scene editing as a next-token prediction task. We leverage a dual-stage training approach combining supervised fine-tuning and preference alignment, enabling a specially trained language model for object addition that accounts for user instructions, spatial geometry, object semantics, and scene-level composition. For scene editing, we employ a zero-shot LLM to handle object removal and prompts for addition. We further introduce a novel voxelization-based evaluation that captures fine-grained geometry beyond 3D bounding boxes. Experimental results surpass state-of-the-art on object addition while maintaining competitive results on full scene synthesis.