DepthEvidence: Unifying Metric Depth Prediction and Geometric Reasoning in Multimodal Language Models

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the disconnect between geometric measurement and linguistic reasoning in multimodal models by proposing DepthEvidence, a 4B-parameter model that enables spatial reasoning under metric constraints. Methodologically, it introduces a novel dense depth-to-language interface employing a camera-conditioned decoder to predict full-resolution metric depth, which is subsequently transformed into object-anchored continuous geometric tokens integrated into the language context. A geometric supervision mechanism is further incorporated to ensure the recoverability of numerical information, alongside the construction of the Depth-VQA benchmark. Experimental results demonstrate that the proposed model achieves state-of-the-art average δ1 performance across nine datasets, rivaling dedicated depth estimators. It excels in instance-level depth estimation and metric reasoning while preserving general visual question answering capabilities.
📝 Abstract
Spatial reasoning with metric constraints requires linking objects to geometric measurements and preserving their numerical content during language reasoning. We present DepthEvidence, a 4B model that uses its own dense metric predictions as object-grounded evidence for language generation. A camera-conditioned decoder predicts full-resolution metric depth using multi-scale visual features and high-resolution RGB refinement. A dense-to-language interface converts predicted depths and decoder features into object-aligned continuous geometry tokens anchored to object identifiers. Geometric supervision encourages metric information to remain recoverable before and after language-context interaction, while instruction tuning supports object measurement and compositional reasoning. We introduce a Depth-VQA benchmark evaluating object-depth queries, relative comparisons, and decisions combining spatial and numerical constraints. Across nine datasets, DepthEvidence achieves the highest average dense $δ_1$ among evaluated methods, competitive with specialized estimators. It also leads the evaluated methods in instance-level metric depth estimation and overall accuracy on both relative and metric reasoning tracks, while broadly preserving general VQA performance and improving spatial understanding relative to the base model.
Problem

Research questions and friction points this paper is trying to address.

metric depth prediction
spatial reasoning
multimodal language models
geometric reasoning
visual question answering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Metric Depth Prediction
Multimodal Language Models
Geometric Reasoning
Dense-to-Language Interface
Depth-VQA