SPOON: Towards Coherent Compositional 3D Scene Generation from Uncalibrated Multi-view Images

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of compositional 3D scene generation from uncalibrated multi-view images, where object pose ambiguity and cross-view inconsistencies lead to disordered scene layouts. To tackle this, the work reformulates the task as geometry-based scene-level pose reasoning and proposes a Guide-Route-Reconcile paradigm. This approach integrates image-conditioned 3D generative priors, multi-view geometric reconstruction, and scene-level pose reasoning algorithms, leveraging reconstructed multi-view geometric cues to iteratively refine object and camera configurations, thereby achieving globally consistent spatial arrangements and eliminating pose ambiguity. Experimental results on the ARSG-110K and MIDI-3D-Front datasets demonstrate substantial improvements in scene consistency, reducing scene-level and object-level Chamfer distances by 12.7% and 17.7%, respectively.
📝 Abstract
Compositional 3D scene generation aims to recover complete 3D object shapes and their spatial arrangement from visual observations. Recent image-conditioned 3D generators provide strong priors for producing high-quality object geometry, making the generation of complex scenes increasingly practical. A central challenge is therefore to spatially organize these generated assets into a globally coherent scene while remaining consistent with multi-view observations. Existing approaches either entangle scene layout with object generation or separately estimate spatial placement from view-specific observations, where pose hypotheses may remain ambiguous and inconsistent across views, often resulting in an incoherent object-camera soup. We introduce SPOON, a framework that reformulates multi-view compositional 3D generation as scene-level, geometry-grounded pose reasoning. Rather than treating view-specific object pose hypotheses independently, SPOON coordinates them using reconstruction-derived multi-view geometry through a Guide-Route-Reconcile paradigm. This progressively organizes object poses and camera configurations into a coherent scene-level spatial arrangement. Extensive experiments on ARSG-110K and MIDI-3D-Front demonstrate consistent improvements in object placement and scene composition across varying numbers of input views. On ARSG-110K, SPOON reduces scene-level and object-level Chamfer distances by 12.7% and 17.7%, respectively, compared with a strong baseline.
Problem

Research questions and friction points this paper is trying to address.

Compositional 3D scene generation
Multi-view images
Scene layout
Pose reasoning
Spatial arrangement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional 3D Scene Generation
Uncalibrated Multi-view Images
Pose Reasoning
Guide-Route-Reconcile
Multi-view Geometry
🔎 Similar Papers
No similar papers found.