GaugeVLM: Structuring Spatial Supervision with Measured Geometric Interventions

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of vision-language models in spatial reasoning, specifically their susceptibility to cross-view contradictions and failure to respond to relational changes. To overcome these issues, we propose GaugeDPO, a method that constructs explicit spatial supervision signals through controlled geometric interventions within 3D scenes. By mapping measurement errors onto preference margins, this approach establishes a joint optimization framework integrating cross-view consistency with intervention constraints. Experimental results demonstrate significant improvements across ten spatial metrics; notably, a 7B-parameter model achieves gains of 15.0 and 18.9 percentage points on distance estimation and spatial relation tasks, respectively. Furthermore, the proposed method successfully generalizes to downstream applications such as autonomous driving, validating its practical efficacy and robustness in real-world scenarios.
📝 Abstract
Vision-language models (VLMs) can contradict themselves across views of the same spatial relation and fail to respond when that relation changes. Addressing these failures requires supervision that captures error magnitude and geometric dependencies across observations, both of which remain implicit in training on individual answers or ordinal preferences. Therefore, we introduce GaugeVLM, which makes this structure explicit through controlled object and camera interventions in explicit 3D scenes, producing linked observations with measured differences between spatial relations and shared truths across views. To translate this structure into learning signals, its core objective, GaugeDPO, converts measured errors into preference margins, directly supervises correct canonical rankings across views, and links intervention-induced answer-odds contrasts to measured relation changes with view-specific scales. Our analysis bounds canonical prediction error and establishes that the cross-view and intervention constraints can be jointly satisfied. Empirically, GaugeVLM improves all 10 established spatial metrics over supervised fine-tuning across three VLM backbones, with the main 7B model gaining 15.0 and 18.9 percentage points on MSMU distance and QSpatial+, respectively. These gains also extend to autonomous driving and embodied reasoning, demonstrating the robust generalization across domains.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Spatial Reasoning
Cross-view Consistency
Geometric Dependencies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Spatial Reasoning
Geometric Interventions
GaugeDPO
3D Scene Supervision
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
H
Hongbo Wang
NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences
Zihan Lin
Zihan Lin
Researcher, Xiaohongshu.
Recommender System
W
Wenkui Yang
NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences; School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences
S
Shiran Ge
NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; National University of Singapore
Yuang Ai
Yuang Ai
MS Student, Institute of Automation, Chinese Academy of Sciences
Computer VisionGenerative ModelsVision-Language Models
Jie Cao
Jie Cao
Institute of Automation, Chinese Academy of Sciences
Computer Vision
Huaibo Huang
Huaibo Huang
NLPR, MAIS, CASIA
Computer VisionGenerative ModelsLow-level VisionFace Recognition
R
Ran He
NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences