🤖 AI Summary
This study addresses the limitation that pretrained robot policies lack explicit metric geometric information, making it difficult to bind reconstructed features to relative spatial representations. To overcome this, we propose a lightweight spatial grafting module that constructs metric grounding tokens and injects them into a flow-matching action expert via cross-attention. This approach seamlessly aligns frozen reconstructed features with robotic relative geometry while preserving pretrained advantages, requiring no modification to the host perception pathway and readily adapting to diverse VLA and WAM models. Experimental results demonstrate that the proposed method achieves a 94% success rate on RoboTwin 2.0, surpassing the strongest baseline, and significantly outperforms the champion solution on long-horizon tasks within the BEHAVIOR-1K benchmark.
📝 Abstract
Pretrained robot manipulation policies such as vision-language-action models (VLAs) or world-action models (WAMs) leave interaction-relevant metric geometry implicit. Recent breakthroughs in spatial reconstruction can supply the necessary geometry reliably, but their features describe local shape without stating where it lies with respect to the robot. How best to deliver these features to a pretrained policy remains unresolved. We propose Spatial Grafting, a versatile, lightweight spatial module that binds frozen reconstruction features to metric, robot-relative geometry. Spatial Grafting constructs metric-grounded spatial tokens and injects them into the flow-matching action expert through cross-attention, without modifying the host's perceptual pathway, so the host retains the full benefit of its pretraining. We evaluate it more broadly than any geometry-aware policy we compare against: one graft architecture, with no per-host redesign, on two VLAs and two WAMs, across four simulation benchmarks that span short-horizon manipulation, visual robustness, clutter and long-horizon mobile manipulation, and on three real-robot platforms with single- and dual-arm configurations. On RoboTwin 2.0, a dual-arm manipulation benchmark, the graft improves every host across VLAs and WAMs. Grafted $\pi_{0.5}$ gains 11.3% and 15.6% on clean and randomized scenes, reaching 94.0% and 92.4%, above the strongest published 3D-conditioned policy, WAM4D (93.8% and 89.9%). The margin widens as the horizon lengthens: on tasks from BEHAVIOR-1K, a dual-arm mobile manipulation challenge scored by average task progress, it surpasses the 2025 challenge winner on five of six tasks,by up to 0.47 Q-score, and exceeds a map-conditioned spatial policy on average across the three tasks both report.