Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited spatial reasoning capabilities of vision-language models arising from the absence of geometric priors in RGB inputs. To this end, it proposes the GPD framework, which injects 3D geometric cues as privileged information into a teacher model for online policy self-distillation while preserving pure RGB inputs during deployment. Core innovations include the first multimodal routing mechanism integrating depth, semantics, and bird’s-eye-view representations, alongside a gated relative policy optimization (GRPO) algorithm that enhances training by imposing privileged KL divergence exclusively on erroneous trajectories. Experimental results demonstrate that this framework substantially improves performance on benchmarks such as VSI-Bench using a 4B-parameter backbone network. It outperforms existing methods without introducing additional inference overhead.
📝 Abstract
Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence. Existing remedies either inject 3D into the model at inference, paying architecture and latency costs, or train with outcome rewards that supervise only the final answer. Spatial errors originate in perception: a misjudged depth or direction can be corrected only by the scene's true geometry, which the 3D-scanned sources of spatial training corpora already provide. We propose GPD (Geometry-Privileged Distillation), which makes geometric evidence the privilege in on-policy self-distillation (OPSD). For each question, depth, semantic, and bird's-eye-view (BEV) cues are rendered as compact text and routed to the teacher alongside the reference answer; a privileged KL, applied only to incorrect trajectories, augments GRPO, and the deployed model remains RGB-only. On the 4B backbone, GPD achieves 57.1 on VSI-Bench and 37.6 average across MindCube, SPARBench, MMSI-Bench, and ViewSpatial, outperforming both GRPO and answer-privileged OPSD across spatial reasoning benchmarks. Ablations confirm the complementarity of 3D and answer privilege, the advantage of question-conditioned routing over full-context injection, and the benefit of restricting distillation to incorrect trajectories.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Spatial Reasoning
3D Geometry
Perception Errors
Innovation

Methods, ideas, or system contributions that make the work stand out.

Geometry-Privileged Distillation
Spatial Reasoning
On-Policy Self-Distillation
GRPO
Vision-Language Models
🔎 Similar Papers
No similar papers found.