Unifying Policy Learning and State Prediction through Spatial Language Modeling

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the decoupling of action policies from state prediction and the lack of geometrically complementary supervision in robotic manipulation by proposing a spatial language modeling framework. This method unifies discrete coordinates and semantic tokens into a shared vocabulary, jointly optimizing action generation and action-conditioned state prediction via an autoregressive Transformer, while introducing shuffle-play pretraining to enhance generalization. Simulation and real-world experiments demonstrate that this single-model architecture significantly outperforms baseline methods in task success rate and target coverage. Furthermore, it accurately captures geometric variations induced by interactions such as pushing, achieving synergistic improvements in both policy learning and state prediction.
📝 Abstract
Learning how actions change scene geometry can provide complementary supervision for goal-directed manipulation. We introduce Spatial Language Modeling, which represents scene contours, goals, action targets, and future states with a shared vocabulary of discrete coordinates and semantic tokens. A task-specific grammar organizes these elements into spatial sequences, allowing one autoregressive Transformer to learn action generation and action-conditioned state prediction through a common next-token objective. We train the model from scratch using random-play transition pretraining followed by joint action and state training on expert demonstrations. During pretraining, recorded action coordinates condition subsequent state predictions and are excluded from the prediction loss. During control, the model decodes only executable action targets and updates its history with newly observed states. We evaluate the approach on Push-T in simulation and on a real robot. The model achieves competitive simulation performance and higher task success and target coverage than the evaluated real-robot policy baselines. Training ablations show improved control with joint action and state sequences, with further gains from random-play pretraining. Given supplied action trajectories, the same model also predicts successive scene states, capturing the geometric effects of pushing.
Problem

Research questions and friction points this paper is trying to address.

policy learning
state prediction
goal-directed manipulation
robot control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Language Modeling
Autoregressive Transformer
Policy Learning
State Prediction
Next-token Objective