SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high inference overhead and external tool dependency of spatially encoded agentic reasoning by proposing an online policy self-distillation framework. The method treats verified agent trajectories as privileged information to internalize spatial reasoning capabilities into a standalone multimodal model. Furthermore, it introduces a novel repetition-aware distillation mechanism that combines repetition masking with unlikelihood regularization to effectively suppress privileged information leakage. Experimental results demonstrate that the proposed model significantly outperforms supervised fine-tuning (SFT) and Group Relative Policy Optimization (GRPO) baselines in average accuracy across both spatial and out-of-distribution benchmarks, exhibiting superior generalization performance.
📝 Abstract
Spatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces. However, this paradigm inherently suffers from prohibitive inference-time overhead and external dependencies. In this paper, we explore whether an MLLM can internalize this agentic capability to operate entirely tool-free. We begin with a simple observation: prompting an MLLM with summarized execution traces of a spatial coding agent naturally unlocks the model's internal spatial Chain-of-Thought (CoT). Motivated by this, we introduce SpatialOPSD, an on-policy self-distillation framework that internalizes spatial reasoning into a standalone MLLM by formulating verified agent traces as privileged information. To mitigate privileged-information leakage during distillation, we introduce Repetition-Aware Distillation, which combines repetition masking with unlikelihood regularization. Experiments across multiple benchmarks demonstrate that self-distilling SpatialOPSD achieves higher average accuracy than SFT and GRPO on both spatial and OOD datasets, exhibiting superior performance and generalization.
Problem

Research questions and friction points this paper is trying to address.

Spatial Reasoning
Multimodal Large Language Models
Self-Distillation
Tool-free Inference
Coding Agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Distillation
Spatial Reasoning
Multimodal Large Language Models
Repetition-Aware Distillation
On-Policy Learning
🔎 Similar Papers
No similar papers found.
R
Rongxue Li
Alibaba Group
M
Meng Yang
Alibaba Group
Y
Yiru Mao
Alibaba Group
Y
Yongliang Tao
Alibaba Group
L
Lulu Hu
Alibaba Group
B
Bin Yang
Alibaba Group
Zhao Xu
Zhao Xu
Principal Staff Engineer, Alibaba Group
Artificial IntelligenceMachine LearningMLSysService DesignCloud Computing
Weihua Luo
Weihua Luo
Alibaba
natural language processingmachine learningartificial intelligence
B
Bowen Xu
Alibaba Group