🤖 AI Summary
This study addresses the systematic failure of vision-language models in video spatial reasoning caused by confounded low-level capabilities, which existing benchmarks struggle to localize. We propose a contrastive diagnostic framework based on shared patterns to decouple perception, geometric representation, and reasoning abilities, enabling precise attribution of four typical error types. Building upon this diagnosis, we construct CROSS, a training-free library of geometric operators, and SpatialClaw, an agent that achieves modular spatial intelligence through typed operators and coordinate contract verification. This approach improves the ReVSI score from 55.9% to 60.2% and attains 66.3% on DSI-Bench, significantly enhancing the reliability of spatial reasoning.
📝 Abstract
Spatial reasoning benchmarks evaluate vision-language models across diverse tasks, but task-level scores do not reveal which underlying capabilities account for success or failure. Each task requires recovering spatial evidence, representing geometry, and reasoning over it. We disentangle these capabilities by comparing predicted and ground-truth spatial context under a shared schema and coordinate contract. This comparison reveals four recurring sources of error: inaccurate perception, missing information in the spatial context, selection of the wrong measurement, and errors in reference frames or in tracking position and orientation. Guided by this diagnosis, we develop CROSS, a training-free library of typed geometric operators and spatial skills that function over available evidence to support reliable video spatial reasoning. The resulting library supplies verified context to non-coding VLMs or callable skills to a SpatialClaw agent. We evaluate \methodname{} on five benchmarks. \methodname{} raises the average score from 55.9\% to 60.2\% on ReVSI and improves the SpatialClaw result from 62.8\% to 66.3\% on DSI-Bench. These gains demonstrate that explicit handling of spatial conventions can repair systematic reasoning failures without additional training.