Score
Designs and implements algorithms and pipelines that, given a single-view image input (including equirectangular panoramas), infer the 3D spatial layout, object instances, and semantic structure of a scene and produce textured, renderable 3D assets and assembled scene representations. This includes methods for predicting room and object geometry, generating instance-level scene variants, creating coherent digital-twin scenes, and producing rendered data for simulation or downstream tasks.
3D scene generation suffers from limited diversity, low visual fidelity, and poor view consistency—hindering its deployment in immersive media, robotics, autonomous driving, and embodied AI. This paper presents a systematic survey of four dominant paradigms: procedural, neural 3D, image-driven, and video-driven generation. We introduce the first unified taxonomy to clarify technical evolution across these approaches. Three emerging frontiers are identified: physics-aware modeling, interactive generation, and perception-generation co-design. Leveraging NeRF, 3D Gaussian Splatting, diffusion models, GANs, and multimodal representations, we conduct a rigorous cross-paradigm evaluation on standard benchmarks, quantifying trade-offs among fidelity, diversity, and view consistency. To foster reproducibility and community advancement, we publicly release an open-source tracking platform that continuously monitors state-of-the-art progress.
研究解决了从单张图片生成可控和可执行3D场景的问题,通过统一视觉-语言-几何框架Fysiverse-3D-Vision实现空间推理与几何重建相互增强。
Existing single-image 3D scene generation methods struggle to simultaneously ensure multi-object geometric consistency and high-fidelity texture synthesis. To address this, we propose a three-stage framework: (1) object decoupling via instance segmentation; (2) pseudo-stereo viewpoint construction and depth inference; and (3) joint optimization of 3D model parameterization and spatial layout to resolve occlusion recovery, camera pose estimation, and 3D–2D alignment. Our approach innovatively integrates image-guided generation, Chamfer distance–minimized layout optimization, and texture-aware inpainting—preserving per-object detail while enhancing inter-object geometric and semantic coherence. Evaluated on multi-object scene benchmarks, our method achieves significant improvements over state-of-the-art: +12.6% in geometric accuracy (CD↓), +9.4% in texture fidelity (LPIPS↓), and +15.2% in layout合理性 (Layout-F1).
This paper introduces a novel paradigm for immersive 3D world generation from a single image—without requiring large-scale training. Addressing the single-image-to-3D-environment reconstruction task, the method proceeds in two stages: first, leveraging a pre-trained diffusion model to synthesize geometrically coherent panoramic images; second, elevating these to 3D via a metric depth estimator and performing 2D inpainting of occluded regions conditioned on rendered point clouds. The core contribution lies in reformulating single-image 3D generation as an in-context learning problem—explicitly modeling 3D structure while bypassing error accumulation inherent in video-synthesis-based approaches. Evaluated on both synthetic and real-world images, the framework produces VR-ready, high-fidelity 3D environments, consistently outperforming state-of-the-art video-synthesis methods across standard metrics including FID, LPIPS, and SSIM.
Manual 3D modeling remains labor-intensive and time-consuming, failing to meet the rapidly growing demands of XR and metaverse applications. Method: This paper presents a systematic survey of state-of-the-art methods for static 3D object and scene generation, introducing— for the first time—a multidimensional classification and cross-comparative framework that jointly considers representation evolution (e.g., point clouds, meshes, NeRFs) and generative paradigms (e.g., supervised learning, diffusion models, 2D foundation model priors, procedural modeling). Contribution/Results: We identify key challenges—including geometric-semantic consistency and scalability—and establish a reusable, multi-axis evaluation framework. Our analysis clarifies technological trajectories and performance boundaries, providing both theoretical foundations and practical guidelines for industrial-grade 3D content generation, thereby advancing the paradigm shift in 3D content creation.
This work proposes a feedforward framework that enables efficient generation of geometrically consistent, full 360° 3D scenes from a single panoramic image—addressing limitations of existing approaches that rely on time-consuming iterative optimization or rigid joint generation strategies and are typically constrained to narrow-field perspective views. By decoupling object generation from layout estimation, the method introduces a plug-and-play Object-World Transformation Predictor and a coarse-to-fine (C2F) alignment mechanism. These components, combined with an enhanced Alignment-VGGT architecture, multi-view rendering, and pseudo-geometric supervision, significantly improve both reconstruction efficiency and accuracy. Experiments demonstrate superior geometric fidelity on both synthetic and real-world datasets, achieving high-fidelity 3D scene generation in approximately 20 seconds on an RTX 4090 GPU.
This work addresses the challenges of single-image 3D scene generation—namely geometric ambiguity, modeling object relationships, and contextual inference—which are exacerbated by the monolithic architectures and heavy reliance on strong supervision in existing methods, limiting their generalization. To overcome these limitations, the authors propose a multi-agent collaborative framework that decouples the generation process into three stages: scene initialization, environment construction, and multi-agent optimization, enabling efficient and precise modeling through structured decomposition. Key innovations include a novel multi-agent mechanism that jointly ensures local refinement and global consistency, and a geometry-aware layout predictor requiring only segmentation-level annotations, substantially reducing dependence on scene-level supervision. Experiments demonstrate that the proposed method significantly outperforms current state-of-the-art approaches across multiple benchmarks, achieving leading performance in geometric accuracy, spatial consistency, and visual realism.
This study addresses the lack of spatial compatibility and physical coherence among objects in single-image 3D scene reconstruction, a challenge particularly pronounced in occluded interaction regions. To overcome this, it introduces the ComOb dataset alongside an explicitly conditioned generative framework grounded in physical relationships. By integrating generative AI, physics simulation, and 3D mesh reconstruction techniques, the proposed method incorporates the geometric and physical relationships of surrounding objects as explicit constraints, enabling collaborative multi-object optimization that restores shape and pose consistency. The approach achieves state-of-the-art performance on both synthetic and real-world scenes, effectively resolving the reconstruction of occluded regions while ensuring that the resulting 3D scenes exhibit both geometric accuracy and physical stability.
This work addresses the challenge of maintaining structural consistency in large-scale 3D scene generation from text, a limitation inherent in current text-to-image/video methods due to their lack of explicit geometric modeling. The authors propose a geometry-first, two-stage framework: first generating an environment-level mesh scaffold—such as walls and floors—from textual descriptions, then conditioning a diffusion model on this scaffold to synthesize photorealistic images while jointly performing semantic segmentation and object reconstruction. By leveraging an explicit geometric backbone, the approach effectively decouples structure from appearance, enabling high object diversity, strong global 3D consistency, and photorealistic detail across arbitrarily scaled multi-room scenes. This facilitates the construction of navigable, realistic 3D environments directly from language prompts.
Generating structurally coherent, asset-disentangled, and textured full 3D indoor scenes from a single 360° equirectangular projection (ERP) image remains highly challenging. This work proposes a three-stage approach: first estimating local geometry as a spatial prior, then introducing a viewpoint-selective cross-attention mechanism to produce a coarse scene layout, and finally synthesizing fine-grained, textured assets through a hybrid global–local attention module combined with a flow-matching strategy. To the best of our knowledge, this is the first method to achieve end-to-end generation of complete 3D indoor scenes from a single 360° image with both structural plausibility and disentangled assets. It significantly outperforms existing approaches across both 2D and 3D evaluation metrics, with its core innovations lying in the designs of viewpoint-selective attention and the global–local hybrid attention mechanism.
This study addresses the challenge of compositional 3D scene generation from uncalibrated multi-view images, where object pose ambiguity and cross-view inconsistencies lead to disordered scene layouts. To tackle this, the work reformulates the task as geometry-based scene-level pose reasoning and proposes a Guide-Route-Reconcile paradigm. This approach integrates image-conditioned 3D generative priors, multi-view geometric reconstruction, and scene-level pose reasoning algorithms, leveraging reconstructed multi-view geometric cues to iteratively refine object and camera configurations, thereby achieving globally consistent spatial arrangements and eliminating pose ambiguity. Experimental results on the ARSG-110K and MIDI-3D-Front datasets demonstrate substantial improvements in scene consistency, reducing scene-level and object-level Chamfer distances by 12.7% and 17.7%, respectively.