🤖 AI Summary
This study addresses the limited reliability of vision-language models (VLMs) in video spatiotemporal reasoning by proposing a training-free "perception-geometry" pipeline. The method leverages deterministic geometric engines—including object segmentation, depth estimation, camera pose recovery, and back-projection—to extract explicit 3D spatial structures from monocular videos and convert them into structured code. This geometric information is directly injected into VLM prompts to compensate for their inherent deficiencies in spatial cognition. Experimental results demonstrate that the proposed approach improves the average accuracy of 2B- and 4B-parameter models by 4.1 points on the VSI-Bench benchmark, with gains reaching up to 24.1 points on absolute distance estimation tasks. These findings indicate that the pipeline significantly enhances the geometric reasoning capabilities of smaller-scale models without requiring additional training.
📝 Abstract
Vision-Language Models (VLMs) are far more reliable at recognizing what appears in a video than at reasoning about its spatial and temporal properties, such as metric distances, object dimensions, and consistent object identities across frames. We present Geometric Code, a perception-to-geometry pipeline that computes explicit spatial structure from video and supplies it to VLMs as context to augment reasoning. A perception layer segments and classifies objects and recovers depth, camera pose, and intrinsics from monocular RGB video. A deterministic geometric engine then back-projects, merges, and cleans these outputs into a spatial code, including per-object positions, dimensions, counts, inter-object distances, appearance order, and room geometry. The code is serialized into VLMs'prompts, either alongside the video or replacing it entirely. Specifically, there is no component trained or fine-tuned in our approach. On VSI-Bench, augmenting 2B and 4B open models with the spatial code improves average accuracy by +4.1 points over the frames-only baseline, with the largest gains on numeric estimation tasks such as absolute distance (+24.1 points). The results suggest that explicitly computed geometry, delivered through the language channel, recovers spatial competence that small VLMs cannot extract from pixels alone.