HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
HARMONY结合代理推理与视觉几何基础,从单幅图像中恢复具有准确物体间关系和高保真度的3D场景。
📝 Abstract
Compositional 3D scene reconstruction has recently been explored from two directions: agentic reasoning that provides semantic understanding of spatial relationships but lacks precise alignment with input images; and visual geometry foundation models that predict dense point maps from input images but the reconstruction quality is limited. Therefore, recovering a complete 3D scene from a single monocular image with accurate inter-object relationships and high-fidelity reconstruction quality remains challenging. In this paper, we present HARMONY, a hierarchical chain-of-thought framework that leverages both agentic reasoning and visual geometry foundation. Given an image of an indoor scene, starting from an empty 3D floorplan, HARMONY first calibrates the camera against the reference image to establish a semantically-grounded spatial frame, then uses agentic VLM reasoning to recover the 3D room layout and an initial placement order. It then places the objects in a hierarchical order, from wall-mounted elements, free-standing furniture, to dependent decorations on top of furniture. We also use depth-first traversal for furniture so each placement conditions on previously resolved structure and a reflective feedback loop to avoid error accumulation. After each object placement by VLM, we use the point cloud estimations to perform geometry-based refinement so that the rendered image aligns better with the input. HARMONY can produce 3D scenes that are semantically consistent and perceptually aligned with the reference image, extending single-image compositional reconstruction to complex indoor scene images. Experiments on synthetic and real-world images demonstrate that HARMONY outperforms the evaluated reconstruction baselines, while qualitative comparisons with GPT-6 Astra suggest more faithful object arrangements and better preservation of scene details.
Problem

Research questions and friction points this paper is trying to address.

monocular image
3D scene reconstruction
agentic reasoning
visual geometry
inter-object relationships
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Agentic Reasoning
Monocular Image-to-Scene Synthesis
Visual Geometry Foundation Models
Depth-first Traversal
Reflective Feedback Loop
🔎 Similar Papers
No similar papers found.