Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the reliance of vision-language models on annotated data for spatial reasoning by proposing an unsupervised self-improvement framework based on self-distillation. The method leverages depth-equivariant geometric priors to construct a privileged teacher model and introduces a recursive training mechanism that maintains teacher stability, thereby enabling the continuous internalization of spatial knowledge. Furthermore, it integrates 3D reconstruction, camera geometry, and dense token-level supervision to facilitate efficient knowledge transfer. Experimental results demonstrate that the proposed approach substantially enhances model performance across multiple benchmarks, achieving state-of-the-art spatial reasoning capabilities among open-source models.
πŸ“ Abstract
Vision-language models (VLMs) increasingly operate in embodied and spatially grounded settings, where accurate understanding of depth, viewpoint, and three-dimensional relations is essential. However, improving spatial reasoning typically relies on ground-truth answers, answer-derived rewards, or other forms of task-specific supervision. We introduce Spatial-OPSD, a label-free self-improvement framework that instead exploits spatial structure naturally available from perception and reconstruction tools. During training, a privileged teacher receives automatically obtainable spatial priors, such as depth, reconstructed 3D relations, and camera geometry, while the student observes only the original visual-language input. On trajectories sampled by the student itself, the teacher provides dense token-level supervision, allowing the student to internalize spatial knowledge without ground-truth answer labels or privileged information at inference time. To extend this supervision beyond a single round, we adopt a round-wise recursive training scheme: the teacher remains frozen within each round to provide a stable learning target, and the improved student initializes both teacher and student in the next round, where privileged spatial priors re-establish an informative teacher--student asymmetry. This enables repeated self-improvement while avoiding a rapidly moving teacher during optimization. Across four VLM families, a single round of Spatial-OPSD consistently improves the five-benchmark average, while three rounds further push a strong spatially specialized model to the open-source frontier, achieving the highest average among the open models and the best results on three of five spatial reasoning benchmarks. Our code is available at https://github.com/vermouth599/Spatial-OPSD.
Problem

Research questions and friction points this paper is trying to address.

spatial reasoning
vision-language models
label-free learning
self-distillation
embodied AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

label-free self-distillation
spatial reasoning
privileged spatial priors
recursive training
vision-language models
πŸ”Ž Similar Papers
No similar papers found.