🤖 AI Summary
This study addresses the lack of spatial compatibility and physical coherence among objects in single-image 3D scene reconstruction, a challenge particularly pronounced in occluded interaction regions. To overcome this, it introduces the ComOb dataset alongside an explicitly conditioned generative framework grounded in physical relationships. By integrating generative AI, physics simulation, and 3D mesh reconstruction techniques, the proposed method incorporates the geometric and physical relationships of surrounding objects as explicit constraints, enabling collaborative multi-object optimization that restores shape and pose consistency. The approach achieves state-of-the-art performance on both synthetic and real-world scenes, effectively resolving the reconstruction of occluded regions while ensuring that the resulting 3D scenes exhibit both geometric accuracy and physical stability.
📝 Abstract
We propose Tetris3D, a generative framework for single-image 3D scene reconstruction that recovers objects which are physically and geometrically coherent as a scene. Existing methods often generate objects independently or couple them implicitly, providing limited guidance for ensuring fine-grained spatial compatibility between neighboring objects that interact with one another. To address this, we explicitly condition the generation of each object on the geometry of surrounding objects and their physical relationships, guiding its shape and pose to remain geometrically and physically plausible within the scene. Moreover, we introduce ComOb, a physics simulation-based dataset of 1.2M scenes featuring physical interactions across diverse object categories, with per-object meshes and pairwise physical relation annotations. Comprehensive experiments on synthetic and realworld scenes show that Tetris3D recovers coherent object shapes and poses even when interacting regions are occluded, and achieves state-of-the-art performance in both generation quality and physical stability.