Spatial Grafting: Grounding 3D Features for Flow-Matching Robot Policies

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that pretrained robot policies lack explicit metric geometric information, making it difficult to bind reconstructed features to relative spatial representations. To overcome this, we propose a lightweight spatial grafting module that constructs metric grounding tokens and injects them into a flow-matching action expert via cross-attention. This approach seamlessly aligns frozen reconstructed features with robotic relative geometry while preserving pretrained advantages, requiring no modification to the host perception pathway and readily adapting to diverse VLA and WAM models. Experimental results demonstrate that the proposed method achieves a 94% success rate on RoboTwin 2.0, surpassing the strongest baseline, and significantly outperforms the champion solution on long-horizon tasks within the BEHAVIOR-1K benchmark.
📝 Abstract
Pretrained robot manipulation policies such as vision-language-action models (VLAs) or world-action models (WAMs) leave interaction-relevant metric geometry implicit. Recent breakthroughs in spatial reconstruction can supply the necessary geometry reliably, but their features describe local shape without stating where it lies with respect to the robot. How best to deliver these features to a pretrained policy remains unresolved. We propose Spatial Grafting, a versatile, lightweight spatial module that binds frozen reconstruction features to metric, robot-relative geometry. Spatial Grafting constructs metric-grounded spatial tokens and injects them into the flow-matching action expert through cross-attention, without modifying the host's perceptual pathway, so the host retains the full benefit of its pretraining. We evaluate it more broadly than any geometry-aware policy we compare against: one graft architecture, with no per-host redesign, on two VLAs and two WAMs, across four simulation benchmarks that span short-horizon manipulation, visual robustness, clutter and long-horizon mobile manipulation, and on three real-robot platforms with single- and dual-arm configurations. On RoboTwin 2.0, a dual-arm manipulation benchmark, the graft improves every host across VLAs and WAMs. Grafted $\pi_{0.5}$ gains 11.3% and 15.6% on clean and randomized scenes, reaching 94.0% and 92.4%, above the strongest published 3D-conditioned policy, WAM4D (93.8% and 89.9%). The margin widens as the horizon lengthens: on tasks from BEHAVIOR-1K, a dual-arm mobile manipulation challenge scored by average task progress, it surpasses the 2025 challenge winner on five of six tasks,by up to 0.47 Q-score, and exceeds a map-conditioned spatial policy on average across the three tasks both report.
Problem

Research questions and friction points this paper is trying to address.

robot manipulation policies
metric geometry
3D features grounding
vision-language-action models
world-action models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Grafting
Flow-Matching
Vision-Language-Action Models
Cross-Attention
3D Geometry Grounding
💼 Related Jobs
No related jobs found.
D
Dingsheng Liu
University of Toronto
Y
Yangzheng Wu
Huawei Noah’s Ark Lab
M
Mahboubeh Asadi
Huawei Noah’s Ark Lab
Z
Zhiyuan Li
University of Toronto
J
Jinbang Huang
Huawei Noah’s Ark Lab
Yixin Xiao
Yixin Xiao
PhD, Ohio State University
GNSS-RGNSS Remote SensingSpectral Method
Tongtong Cao
Tongtong Cao
Researcher, Huawei Noah's Ark Lab
RoboticsEmbodied AIAutonomous driving
Yingxue Zhang
Yingxue Zhang
Huawei
Graph representation learningGraph ReasoningLLMs ReasoningKnowledge GraphsRecommender Systems