SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of existing benchmarks that evaluate only static spatial perception while neglecting predictive spatial reasoning. We introduce the first benchmark designed to directly diagnose this capability through an "observe-transform-infer" framework, which decomposes tasks into three hierarchical levels: static perception, local prediction, and global prediction. Leveraging procedural generation from real-world scenes, we design sixteen task categories and conduct evaluations incorporating bridging views and explicit 3D evidence. Experimental results demonstrate that even state-of-the-art models perform significantly below human baselines, validating the critical role of bridging views. Furthermore, fine-tuning on our proposed dataset improves the accuracy of Qwen3-VL-4B to 65.7%.
📝 Abstract
Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input. Yet real-world spatial intelligence demands predictive spatial reasoning: constructing a scene from observations, anticipating how an intervention changes it, and reasoning about the unseen outcome. We introduce SpaceCast-Bench, the first benchmark to directly and diagnostically evaluate this capability. Built around an observe-transform-infer framework, its 3,862 questions from 182 real-world scenes span 16 task types at three levels: static perception, local prediction, and global prediction, progressively requiring scene understanding, spatial state updating, and relational inference over unobserved outcomes. Evaluating 21 models exposes a stark gap: the strongest model reaches only 58.0% against 87.2% human performance, while spatially specialized models remain near random chance. Controlled analyses further reveal that bridge views are critical for integrating distributed observations, and that explicit 3D evidence benefits models more reliably than generated outcome images or videos. Fine-tuning on our programmatically generated data lifts Qwen3-VL-4B from 34.0% to 65.7% with macro-average gains across six out-of-domain benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Predictive Spatial Reasoning
Vision-Language Models
Spatial Perception
Benchmark Evaluation
Scene Understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Predictive Spatial Reasoning
Vision-Language Models
Benchmark
Observe-Transform-Infer
3D Evidence
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hongxing Li
Zhejiang University
J
Jinyue Su
Zhejiang University
D
Dingming Li
Zhejiang University
Wenqi Zhang
Wenqi Zhang
Zhejiang University
Language ModelMultimodal LearningEmbodied Agents
Weiming Lu
Weiming Lu
Zhejiang University
Natural Language ProcessingLarge Language ModelsAGI
J
Jun Xiao
Zhejiang University
Y
Yueting Zhuang
Zhejiang University
Y
Yongliang Shen
Zhejiang University