single-view 3d layout inference

Designs and implements algorithms and pipelines that, given a single-view image input (including equirectangular panoramas), infer the 3D spatial layout, object instances, and semantic structure of a scene and produce textured, renderable 3D assets and assembled scene representations. This includes methods for predicting room and object geometry, generating instance-level scene variants, creating coherent digital-twin scenes, and producing rendered data for simulation or downstream tasks.

single-view3dlayoutinference

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Towards Geometric and Textural Consistency 3D Scene Generation via Single Image-guided Model Generation and Layout Optimization

Jul 20, 2025
XT
Xiang Tang
🏛️ Harbin Institute of Technology | Peng Cheng Laboratory | Harbin Institute of Technology, Suzhou Research Institute

Existing single-image 3D scene generation methods struggle to simultaneously ensure multi-object geometric consistency and high-fidelity texture synthesis. To address this, we propose a three-stage framework: (1) object decoupling via instance segmentation; (2) pseudo-stereo viewpoint construction and depth inference; and (3) joint optimization of 3D model parameterization and spatial layout to resolve occlusion recovery, camera pose estimation, and 3D–2D alignment. Our approach innovatively integrates image-guided generation, Chamfer distance–minimized layout optimization, and texture-aware inpainting—preserving per-object detail while enhancing inter-object geometric and semantic coherence. Evaluated on multi-object scene benchmarks, our method achieves significant improvements over state-of-the-art: +12.6% in geometric accuracy (CD↓), +9.4% in texture fidelity (LPIPS↓), and +15.2% in layout合理性 (Layout-F1).

Ensuring geometric and textural consistency in multi-object 3D generationGenerating 3D scenes from single RGB images with quality and coherenceOptimizing layout parameters for precise alignment with input guidance

A Recipe for Generating 3D Worlds From a Single Image

Mar 20, 2025
KS
Katja Schwarz
🏛️ Meta Reality Labs

This paper introduces a novel paradigm for immersive 3D world generation from a single image—without requiring large-scale training. Addressing the single-image-to-3D-environment reconstruction task, the method proceeds in two stages: first, leveraging a pre-trained diffusion model to synthesize geometrically coherent panoramic images; second, elevating these to 3D via a metric depth estimator and performing 2D inpainting of occluded regions conditioned on rendered point clouds. The core contribution lies in reformulating single-image 3D generation as an in-context learning problem—explicitly modeling 3D structure while bypassing error accumulation inherent in video-synthesis-based approaches. Evaluated on both synthetic and real-world images, the framework produces VR-ready, high-fidelity 3D environments, consistently outperforming state-of-the-art video-synthesis methods across standard metrics including FID, LPIPS, and SSIM.

Generating 3D worlds from single imagesOutperforming video synthesis-based methods in qualityUsing 2D inpainting models with minimal training

Recent Advance in 3D Object and Scene Generation: A Survey

Apr 16, 2025
XT
Xiang Tang
🏛️ Harbin Institute of Technology | Pengcheng Laboratory

Manual 3D modeling remains labor-intensive and time-consuming, failing to meet the rapidly growing demands of XR and metaverse applications. Method: This paper presents a systematic survey of state-of-the-art methods for static 3D object and scene generation, introducing— for the first time—a multidimensional classification and cross-comparative framework that jointly considers representation evolution (e.g., point clouds, meshes, NeRFs) and generative paradigms (e.g., supervised learning, diffusion models, 2D foundation model priors, procedural modeling). Contribution/Results: We identify key challenges—including geometric-semantic consistency and scalability—and establish a reusable, multi-axis evaluation framework. Our analysis clarifies technological trajectories and performance boundaries, providing both theoretical foundations and practical guidelines for industrial-grade 3D content generation, thereby advancing the paradigm shift in 3D content creation.

Addressing challenges in 3D content creation for XR/MetaverseOvercoming labor-intensive manual 3D modeling limitationsSurveying AI-driven 3D object and scene generation methods

This work proposes a feedforward framework that enables efficient generation of geometrically consistent, full 360° 3D scenes from a single panoramic image—addressing limitations of existing approaches that rely on time-consuming iterative optimization or rigid joint generation strategies and are typically constrained to narrow-field perspective views. By decoupling object generation from layout estimation, the method introduces a plug-and-play Object-World Transformation Predictor and a coarse-to-fine (C2F) alignment mechanism. These components, combined with an enhanced Alignment-VGGT architecture, multi-view rendering, and pseudo-geometric supervision, significantly improve both reconstruction efficiency and accuracy. Experiments demonstrate superior geometric fidelity on both synthetic and real-world datasets, achieving high-fidelity 3D scene generation in approximately 20 seconds on an RTX 4090 GPU.

360-degree environmentcompositional 3D scene generationimage-to-3D

Latest Papers

What's happening recently
View more

This work addresses the challenges of single-image 3D scene generation—namely geometric ambiguity, modeling object relationships, and contextual inference—which are exacerbated by the monolithic architectures and heavy reliance on strong supervision in existing methods, limiting their generalization. To overcome these limitations, the authors propose a multi-agent collaborative framework that decouples the generation process into three stages: scene initialization, environment construction, and multi-agent optimization, enabling efficient and precise modeling through structured decomposition. Key innovations include a novel multi-agent mechanism that jointly ensures local refinement and global consistency, and a geometry-aware layout predictor requiring only segmentation-level annotations, substantially reducing dependence on scene-level supervision. Experiments demonstrate that the proposed method significantly outperforms current state-of-the-art approaches across multiple benchmarks, achieving leading performance in geometric accuracy, spatial consistency, and visual realism.

3D scene generationenvironmental contextgeometric consistency

This study addresses the lack of spatial compatibility and physical coherence among objects in single-image 3D scene reconstruction, a challenge particularly pronounced in occluded interaction regions. To overcome this, it introduces the ComOb dataset alongside an explicitly conditioned generative framework grounded in physical relationships. By integrating generative AI, physics simulation, and 3D mesh reconstruction techniques, the proposed method incorporates the geometric and physical relationships of surrounding objects as explicit constraints, enabling collaborative multi-object optimization that restores shape and pose consistency. The approach achieves state-of-the-art performance on both synthetic and real-world scenes, effectively resolving the reconstruction of occluded regions while ensuring that the resulting 3D scenes exhibit both geometric accuracy and physical stability.

3D scene reconstructiongeometric plausibilityphysical coherence

This work addresses the challenge of maintaining structural consistency in large-scale 3D scene generation from text, a limitation inherent in current text-to-image/video methods due to their lack of explicit geometric modeling. The authors propose a geometry-first, two-stage framework: first generating an environment-level mesh scaffold—such as walls and floors—from textual descriptions, then conditioning a diffusion model on this scaffold to synthesize photorealistic images while jointly performing semantic segmentation and object reconstruction. By leveraging an explicit geometric backbone, the approach effectively decouples structure from appearance, enabling high object diversity, strong global 3D consistency, and photorealistic detail across arbitrarily scaled multi-room scenes. This facilitates the construction of navigable, realistic 3D environments directly from language prompts.

3D scene generationgeometry representationlarge-scale environments

Generating structurally coherent, asset-disentangled, and textured full 3D indoor scenes from a single 360° equirectangular projection (ERP) image remains highly challenging. This work proposes a three-stage approach: first estimating local geometry as a spatial prior, then introducing a viewpoint-selective cross-attention mechanism to produce a coarse scene layout, and finally synthesizing fine-grained, textured assets through a hybrid global–local attention module combined with a flow-matching strategy. To the best of our knowledge, this is the first method to achieve end-to-end generation of complete 3D indoor scenes from a single 360° image with both structural plausibility and disentangled assets. It significantly outperforms existing approaches across both 2D and 3D evaluation metrics, with its core innovations lying in the designs of viewpoint-selective attention and the global–local hybrid attention mechanism.

360° image3D indoor scene generationscene layout

This study addresses the challenge of compositional 3D scene generation from uncalibrated multi-view images, where object pose ambiguity and cross-view inconsistencies lead to disordered scene layouts. To tackle this, the work reformulates the task as geometry-based scene-level pose reasoning and proposes a Guide-Route-Reconcile paradigm. This approach integrates image-conditioned 3D generative priors, multi-view geometric reconstruction, and scene-level pose reasoning algorithms, leveraging reconstructed multi-view geometric cues to iteratively refine object and camera configurations, thereby achieving globally consistent spatial arrangements and eliminating pose ambiguity. Experimental results on the ARSG-110K and MIDI-3D-Front datasets demonstrate substantial improvements in scene consistency, reducing scene-level and object-level Chamfer distances by 12.7% and 17.7%, respectively.

Compositional 3D scene generationMulti-view imagesPose reasoning

Hot Scholars

MR

Michael R. Lyu

Professor of Computer Science & Engineering, The Chinese University of Hong Kong
software engineeringsoftware reliabilityfault tolerancemachine learning
FL

Fang-Lue Zhang

Senior Lecturer (Associate Professor), Victoria University of Wellington, New Zealand
Computer GraphicsImage and Video ProcessingVR AR MRComputer Vision
SL

Shuqing Li

The Chinese University of Hong Kong
Reliable Spatial IntelligenceMultimodal LLM AgentsXR (VR/AR/MR) SystemXR Security
XJ

Xu Jia

Associate Professor at Dalian University of Technology
Computer VisionMachine LearningBio-Inspired Vision
DB

Davide Boscaini

Fondazione Bruno Kessler
Geometric Deep LearningComputer Vision