🤖 AI Summary
This study addresses the spatial misalignment arising from the decoupling of monocular depth estimation and 3D visual grounding tasks by proposing a unified agent framework based on large language models (LLMs). The method leverages LLMs for spatial-visual program planning, orchestrating vision-language models, object detectors, and 3D back-projection techniques to enable flexible multi-task reasoning for instance-level depth estimation and 3D bounding box prediction. Furthermore, an M3SI benchmark dataset is constructed to evaluate monocular 3D spatial understanding capabilities. Experimental results demonstrate that the proposed approach surpasses existing state-of-the-art models in both instance-level depth estimation, achieving 52.61% at δ<0.25, and 3D grounding, attaining a mean Intersection over Union (mIoU) of 41.73%.
📝 Abstract
Monocular metric depth estimation and 3D visual grounding represent the two complementary cornerstones of monocular 3D spatial understanding (M3Sun), from which the fundamental 3D spatial information required by M3Sun can be acquired. However, these complementary tasks are generally conducted by separate frameworks, which pose challenges of inflexible and unaligned spatial information access for embodied intelligence systems. In this paper, we propose a unified agent for monocular 3D spatial understanding (M3SunAgent) that leverages a large language model (LLM) as a task planner for spatial visual programming, which flexibly generate structured programs and coordinate tools. For instance-level metric depth estimation task, M3SunAgent invokes an object detector tool to locate the target, estimates depth at selected points with a depth estimation tool, and aggregates these predictions into an instance-level depth estimate. We also construct the M3Sun Instance (M3SI) dataset, a benchmark with 2,910 samples for evaluation. For monocular 3D visual grounding task, M3SunAgent uses a vision-language model (VLM) tool to locate the target and output basic spatial attributes, then combines back-projection tool with a dimension-lifting tool to predict its 3D bounding box. Experimental results demonstrate the superior performance of M3SunAgent. Specifically, in evaluations of instance-level monocular metric depth estimation, M3SunAgent achieves the best performance among all compared models, 52.61% of predicted instances are distributed below depth error 0.25 ($δ< 0.25$). In evaluations of monocular 3D visual grounding, M3SunAgent demonstrates overall competitive performance than vision and VLM models, reaching a 3D mean intersection over union (mIoU) of 41.73% and exceeding the state-of-the-art MonoVLM model by 3.62%.