M3SunAgent: Monocular 3D Spatial Understanding Agent for Metric Depth Estimation and 3D Visual Grounding

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the spatial misalignment arising from the decoupling of monocular depth estimation and 3D visual grounding tasks by proposing a unified agent framework based on large language models (LLMs). The method leverages LLMs for spatial-visual program planning, orchestrating vision-language models, object detectors, and 3D back-projection techniques to enable flexible multi-task reasoning for instance-level depth estimation and 3D bounding box prediction. Furthermore, an M3SI benchmark dataset is constructed to evaluate monocular 3D spatial understanding capabilities. Experimental results demonstrate that the proposed approach surpasses existing state-of-the-art models in both instance-level depth estimation, achieving 52.61% at δ<0.25, and 3D grounding, attaining a mean Intersection over Union (mIoU) of 41.73%.
📝 Abstract
Monocular metric depth estimation and 3D visual grounding represent the two complementary cornerstones of monocular 3D spatial understanding (M3Sun), from which the fundamental 3D spatial information required by M3Sun can be acquired. However, these complementary tasks are generally conducted by separate frameworks, which pose challenges of inflexible and unaligned spatial information access for embodied intelligence systems. In this paper, we propose a unified agent for monocular 3D spatial understanding (M3SunAgent) that leverages a large language model (LLM) as a task planner for spatial visual programming, which flexibly generate structured programs and coordinate tools. For instance-level metric depth estimation task, M3SunAgent invokes an object detector tool to locate the target, estimates depth at selected points with a depth estimation tool, and aggregates these predictions into an instance-level depth estimate. We also construct the M3Sun Instance (M3SI) dataset, a benchmark with 2,910 samples for evaluation. For monocular 3D visual grounding task, M3SunAgent uses a vision-language model (VLM) tool to locate the target and output basic spatial attributes, then combines back-projection tool with a dimension-lifting tool to predict its 3D bounding box. Experimental results demonstrate the superior performance of M3SunAgent. Specifically, in evaluations of instance-level monocular metric depth estimation, M3SunAgent achieves the best performance among all compared models, 52.61% of predicted instances are distributed below depth error 0.25 ($δ< 0.25$). In evaluations of monocular 3D visual grounding, M3SunAgent demonstrates overall competitive performance than vision and VLM models, reaching a 3D mean intersection over union (mIoU) of 41.73% and exceeding the state-of-the-art MonoVLM model by 3.62%.
Problem

Research questions and friction points this paper is trying to address.

monocular 3D spatial understanding
metric depth estimation
3D visual grounding
embodied intelligence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Monocular 3D Spatial Understanding
Large Language Model
Visual Programming
Metric Depth Estimation
3D Visual Grounding
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Jinsong Zhang
Jinsong Zhang
Université Laval
Computer VisionDeep LearningComputer Graphics
K
Kejun Wu
School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China
M
Ming Zhu
School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China
R
Renjie Qiao
College of Intelligent Systems Science and Engineering, Harbin Engineering University, Harbin 150001, China
C
Chengtao Cai
College of Intelligent Systems Science and Engineering, Harbin Engineering University, Harbin 150001, China
Zhengguo Li
Zhengguo Li
IEEE Fellow, Senior Principal Scientist, Institute for Infocomm Research
Video codingPhysics-guided AIComputational photographySensor fusionSwitched control