SpatialSkill: Self-Evolving Skills for Cross-View Spatial Reasoning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that visual-language models encode spatial reasoning knowledge implicitly within their weights during cross-view spatial reasoning, hindering transferability and auditability. To overcome this, we propose a weight-update-free framework enabling frozen models to extract explicit natural language reasoning skills from offline trajectories. We introduce a novel skill filtering mechanism grounded in visual grounding and executability verification, employ classification-based routing to mitigate negative transfer, and construct versioned skill manuals for multi-stage validation. Evaluated on the CityCube benchmark, our approach yields substantial performance improvements, with a 9B open-source model surpassing the strongest closed-source baselines while maintaining transparent and auditable reasoning strategies.
📝 Abstract
Cross-view spatial reasoning requires a model to align different viewpoints into a coherent spatial representation, yet this ability remains challenging for vision-language models despite being natural to humans. Existing methods typically improve spatial reasoning by updating model weights, which keeps the acquired knowledge implicit and tied to a specific backbone. We propose \textit{SpatialSkill}, a weight-update-free framework that enables a frozen vision-language model to accumulate explicit natural-language reasoning skills from offline trajectories. Unlike symbolic tasks, perceptual skills cannot be reliably verified simply by executing them: a plausible spatial rule may lack visual support or require transformations that the frozen model cannot perform. SpatialSkill therefore admits candidate skills only after visual-grounding and executability checks, constrains manual evolution to prevent harmful regressions, and routes skills by spatial-reasoning category to reduce negative transfer. On CityCube, across four frozen executors, SpatialSkill yields consistent gains, and a 9B executor equipped with SpatialSkill surpasses the strongest closed-source reference in our evaluation. The skills are stored in a versioned natural-language manual, making the reasoning strategies explicit and auditable without modifying model parameters. Code at https://github.com/vindahi/SpatialSkill.
Problem

Research questions and friction points this paper is trying to address.

cross-view spatial reasoning
vision-language models
weight-update-free
perceptual skills
visual grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-View Spatial Reasoning
Weight-Free Framework
Self-Evolving Skills
Vision-Language Models
Visual Grounding
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ruifan Zuo
Shandong University
G
Guocheng Hu
Shandong University
W
Wanshui Gan
Shanghai AI Laboratory
Junyi Wang
Junyi Wang
University of Electronic Science and Tenchonolegy of China
Image RegistrationMRI
X
Xiang Lei
Zhiyang Innovation Co., Ltd.
T
Tian Gan
Shandong University