SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited spatial chain-of-thought (CoT) performance of vision-language models in multi-view spatial reasoning caused by the absence of geometric intermediate supervision. To this end, we propose a two-stage framework. First, we introduce QA-native reconstruction pretraining, which reformulates geometric estimations—such as 3D points and object-centric queries—into textual question-answering, enabling local details and global context to share a unified autoregressive interface. Subsequently, CoT training with a visual compensation mechanism is incorporated to optimize answer derivation. This approach increases CoT gains on the ReVSI benchmark from 2.6 to 6.9 points and achieves state-of-the-art performance on both ReVSI and VSI-Bench, surpassing the strongest baseline by 8.7 points.
📝 Abstract
Vision-language models (VLMs) can benefit from geometric priors for multi-view spatial reasoning, yet answer-only training does not directly supervise the intermediate geometric estimates and their use in deriving quantitative spatial answers. We hypothesize that spatial chain-of-thought (CoT) supervision becomes more effective when the VLM first jointly learns complementary local geometry and global scene context through multi-view reconstruction. We introduce SpatialSpeak, a two-stage framework that connects QA-native reconstruction pretraining with spatial CoT learning. In Stage I, QA-Native Reconstruction Pretraining (QA-RP) combines marked-point 3D queries for fine-grained local geometry with object-center queries for global scene context across views. Both tasks are formulated as text-based question answering, allowing geometric estimation and subsequent reasoning to share the same autoregressive output interface. In Stage II, spatial CoT with Visual Compensation (CoT-VC) trains the model to express question-relevant geometric estimates and use them to derive answers, with reliability assessment and visual compensation supporting answer refinement when needed. On ReVSI, QA-RP increases the gain from CoT-VC from 2.6 to 6.9 points, and ablations show that both local and global reconstruction supervision are beneficial. SpatialSpeak achieves state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench, with a ReVSI score of 62.8 that exceeds the strongest compared baseline by 8.7 points.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Spatial Reasoning
Chain-of-Thought
Multi-view Reconstruction
Geometric Priors
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Chain-of-Thought
QA-Native Reconstruction
Vision-Language Models
Multi-view Spatial Reasoning
Visual Compensation
🔎 Similar Papers
No similar papers found.