🤖 AI Summary
This work addresses key challenges in 3D indoor scene generation from multimodal inputs (images, sketches, or text), including geometric inconsistency, frequent spatial conflicts, and incomplete layout synthesis. We propose a multi-stage generative framework that jointly incorporates semantic structure and spatial constraints. Our approach introduces a unified constraint representation and a conflict-aware localization strategy, integrates heuristic depth-first search (HDFS) to optimize object placement order, and designs a constraint-driven layout refinement module. To enhance semantic alignment, we incorporate natural language understanding into the generation pipeline. Experimental results demonstrate that our method consistently outperforms state-of-the-art approaches across all input modalities: furniture collision rate is reduced by 32.7%, and layout completeness improves by 28.4%. The generated scenes exhibit significantly enhanced spatial plausibility, visual realism, and semantic coherence.
📝 Abstract
Generating realistic 3D indoor scenes from user inputs remains a challenging problem in computer vision and graphics, requiring careful balance of geometric consistency, spatial relationships, and visual realism. While neural generation methods often produce repetitive elements due to limited global spatial reasoning, procedural approaches can leverage constraints for controllable generation but struggle with multi-constraint scenarios. When constraints become numerous, object collisions frequently occur, forcing the removal of furniture items and compromising layout completeness.
To address these limitations, we propose RoomCraft, a multi-stage pipeline that converts real images, sketches, or text descriptions into coherent 3D indoor scenes. Our approach combines a scene generation pipeline with a constraint-driven optimization framework. The pipeline first extracts high-level scene information from user inputs and organizes it into a structured format containing room type, furniture items, and spatial relations. It then constructs a spatial relationship network to represent furniture arrangements and generates an optimized placement sequence using a heuristic-based depth-first search (HDFS) algorithm to ensure layout coherence. To handle complex multi-constraint scenarios, we introduce a unified constraint representation that processes both formal specifications and natural language inputs, enabling flexible constraint-oriented adjustments through a comprehensive action space design. Additionally, we propose a Conflict-Aware Positioning Strategy (CAPS) that dynamically adjusts placement weights to minimize furniture collisions and ensure layout completeness.
Extensive experiments demonstrate that RoomCraft significantly outperforms existing methods in generating realistic, semantically coherent, and visually appealing room layouts across diverse input modalities.