From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing vision-language models (VLMs) in predictive reasoning over unobserved future spatial states. To this end, we propose SpatialMind, a model integrated with a metric-level spatial reasoning framework. Specifically, it employs a metric depth adapter to extract precise spatial representations and leverages a progressive state chain to enable temporal deduction from current visual observations to future spatial relationships. Furthermore, an automated question-answering generation technique is introduced to construct a scalable data engine, alongside a three-tiered predictive benchmark established for systematic evaluation. Experimental results demonstrate that the proposed approach significantly outperforms state-of-the-art methods on our custom benchmark while maintaining strong competitiveness across multiple external datasets, thereby offering a novel paradigm for spatial predictive reasoning in VLMs.
📝 Abstract
Predicting future spatial states supports collision avoidance and timely decision-making in dynamic environments. However, existing vision-language models (VLMs) and benchmarks for spatial reasoning primarily focus on observed scenes, leaving predictive spatial reasoning beyond the observed interval underexplored. To this end, we introduce SpatialMind, a metric-scale VLM for spatial reasoning and future prediction. Its metric depth adapter anchors spatial reasoning to real-world scale, while its progressive state chain establishes current spatial states and observed dynamics as the foundation for future prediction. Given a video prefix, SpatialMind predicts distances, motion directions, and spatial relations in both observed and unseen future frames. For training and evaluation, we build a scalable data engine that grounds entity descriptions in metric geometry to generate question-answer pairs and state supervision. Using this engine, we construct the SpatialMind-30K dataset and the SpatialMind-2K benchmark, both covering driving and everyday egocentric scenes. The benchmark spans eight tasks across three levels: current-state understanding, observed-dynamics understanding, and future prediction. Experiments show that SpatialMind substantially outperforms both general and spatially specialized models on our benchmark while achieving competitive zero-shot performance on VSI-Bench, OSI-Bench, and VLM4D.
Problem

Research questions and friction points this paper is trying to address.

Predictive Spatial Reasoning
Vision-Language Models
Future Prediction
Dynamic Environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

Predictive Spatial Reasoning
Metric Depth Adapter
Progressive State Chain
Vision-Language Models
Scalable Data Engine
🔎 Similar Papers
No similar papers found.