🤖 AI Summary
This study addresses the challenge of generating 3D scenes from a single image, where the lack of geometric context between objects and their environments leads to poor spatial coherence. To overcome this, the authors propose a joint generation framework based on a shared coordinate system that enables unified modeling of both environments and objects. The approach introduces two core innovations: a scene-frame generation mechanism coupled with object-centric local-frame refinement, which facilitates explicit environmental modeling and effective transfer of pretrained priors; and an object-to-scene knowledge distillation strategy leveraging automatically synthesized rendering data. Extensive evaluations on indoor and outdoor benchmarks demonstrate that the proposed method significantly improves scene-level spatial coherence, consistently outperforming existing baselines across all metrics.
📝 Abstract
We present DistScene, a framework for single-image compositional 3D scene generation by jointly modeling the environment and individual objects. Unlike existing methods that represent scenes primarily as collections of objects, we model the environment as an explicit scene component to provide geometric context for object placement. Specifically, we introduce Scene-Frame Generation, which jointly generates separate environment and object components in a shared coordinate frame, allowing their geometry and relative placement to be learned together. Then we introduce Object-Centric Refinement to refine each object in a local frame with scene context. Finally, we develop Object-to-Scene Distillation to transfer pretrained object-generation priors to scene generation through automatically composed and rendered synthetic scenes. Evaluations on indoor and outdoor benchmarks demonstrate improved scene-level spatial coherence over the evaluated baselines. Project page: https://coolbeam.github.io/DistScene/