🤖 AI Summary
This study addresses the challenge of generating complete 3D scenes from a single image, where existing methods struggle to accommodate both indoor and outdoor environments while reconstructing geometry in unobserved regions. Building upon the Trellis 2 architecture, this work reformulates an object-centric generator by introducing a distance-aware adaptive chunking strategy to enhance the perception of free space and unobserved areas. Furthermore, explicit 2D-3D correspondence and feature lifting techniques are incorporated to ensure geometric consistency. To facilitate training, a large-scale synthetic outdoor dataset is constructed. Experimental results demonstrate that the proposed method significantly outperforms existing baselines across multiple benchmarks in terms of both geometric accuracy and perceptual quality, achieving high-fidelity 3D mesh generation for diverse indoor and outdoor scenes.
📝 Abstract
Single-image scene generation aims to produce a complete 3D scene mesh from a single image, including surfaces the camera did not observe. While pretrained 3D object generators encode a strong shape prior, they are mainly designed for isolated objects in a fixed canonical volume and focus mostly on indoor scenes, since diverse 3D data for outdoor scenes are quite limited. In this work, we present a method that redesigns such an object-centric generator, e.g., Trellis 2, to work on both indoor and outdoor scenes while retaining its prior. We accomplish this by (a) partitioning the scene into adaptive chunks that scale relative to the distance to the camera; nearby chunks have a smaller size to keep the finer detail, while distant structures, e.g., buildings, are covered by large chunks; (b) making the generator capture explicit 2D-3D correspondence by lifting image features and making the model aware of the free space, observed surface, and unobserved region; (c) synthesizing around 4,000 outdoor scenes to broaden the training data, as existing scene datasets are largely indoor. Experiments on Tanks and Temples, ScanNet++, and in-the-wild images show that our method outperforms all baselines in geometric accuracy and perceptual quality across both indoor and outdoor scenes.