Geometric Encoding for Spatial Reasoning in Vision-Language Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited reliability of vision-language models (VLMs) in video spatiotemporal reasoning by proposing a training-free "perception-geometry" pipeline. The method leverages deterministic geometric engines—including object segmentation, depth estimation, camera pose recovery, and back-projection—to extract explicit 3D spatial structures from monocular videos and convert them into structured code. This geometric information is directly injected into VLM prompts to compensate for their inherent deficiencies in spatial cognition. Experimental results demonstrate that the proposed approach improves the average accuracy of 2B- and 4B-parameter models by 4.1 points on the VSI-Bench benchmark, with gains reaching up to 24.1 points on absolute distance estimation tasks. These findings indicate that the pipeline significantly enhances the geometric reasoning capabilities of smaller-scale models without requiring additional training.
📝 Abstract
Vision-Language Models (VLMs) are far more reliable at recognizing what appears in a video than at reasoning about its spatial and temporal properties, such as metric distances, object dimensions, and consistent object identities across frames. We present Geometric Code, a perception-to-geometry pipeline that computes explicit spatial structure from video and supplies it to VLMs as context to augment reasoning. A perception layer segments and classifies objects and recovers depth, camera pose, and intrinsics from monocular RGB video. A deterministic geometric engine then back-projects, merges, and cleans these outputs into a spatial code, including per-object positions, dimensions, counts, inter-object distances, appearance order, and room geometry. The code is serialized into VLMs'prompts, either alongside the video or replacing it entirely. Specifically, there is no component trained or fine-tuned in our approach. On VSI-Bench, augmenting 2B and 4B open models with the spatial code improves average accuracy by +4.1 points over the frames-only baseline, with the largest gains on numeric estimation tasks such as absolute distance (+24.1 points). The results suggest that explicitly computed geometry, delivered through the language channel, recovers spatial competence that small VLMs cannot extract from pixels alone.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Spatial Reasoning
Geometric Encoding
Video Understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Geometric Encoding
Spatial Reasoning
Vision-Language Models
Training-free Pipeline
Monocular 3D Reconstruction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Antonio Jun
Dept. of Computer Science, Hunter College, New York City, United States
H
Haoshui Yu
School of Arts and Sciences, New York University, New York City, United States
Z
Zhengyi Lu
Dept. of Engineering and Computer Science, Oakland University, Rochester, United States
Huirong Fu
Huirong Fu
Oakland University
NetworksSecurityPrivacyTrust ManagementInternet Data Center
Yao Qiang
Yao Qiang
Oakland University
Trustworthy AINatural Language ProcessingLarge Language ModelMachine Learning