🤖 AI Summary
This study addresses the disconnect between geometric measurement and linguistic reasoning in multimodal models by proposing DepthEvidence, a 4B-parameter model that enables spatial reasoning under metric constraints. Methodologically, it introduces a novel dense depth-to-language interface employing a camera-conditioned decoder to predict full-resolution metric depth, which is subsequently transformed into object-anchored continuous geometric tokens integrated into the language context. A geometric supervision mechanism is further incorporated to ensure the recoverability of numerical information, alongside the construction of the Depth-VQA benchmark. Experimental results demonstrate that the proposed model achieves state-of-the-art average δ1 performance across nine datasets, rivaling dedicated depth estimators. It excels in instance-level depth estimation and metric reasoning while preserving general visual question answering capabilities.
📝 Abstract
Spatial reasoning with metric constraints requires linking objects to geometric measurements and preserving their numerical content during language reasoning. We present DepthEvidence, a 4B model that uses its own dense metric predictions as object-grounded evidence for language generation. A camera-conditioned decoder predicts full-resolution metric depth using multi-scale visual features and high-resolution RGB refinement. A dense-to-language interface converts predicted depths and decoder features into object-aligned continuous geometry tokens anchored to object identifiers. Geometric supervision encourages metric information to remain recoverable before and after language-context interaction, while instruction tuning supports object measurement and compositional reasoning. We introduce a Depth-VQA benchmark evaluating object-depth queries, relative comparisons, and decisions combining spatial and numerical constraints. Across nine datasets, DepthEvidence achieves the highest average dense $δ_1$ among evaluated methods, competitive with specialized estimators. It also leads the evaluated methods in instance-level metric depth estimation and overall accuracy on both relative and metric reasoning tracks, while broadly preserving general VQA performance and improving spatial understanding relative to the base model.