🤖 AI Summary
This study addresses the failure of multimodal foundation models in fine-grained robotic manipulation caused by the limited observational capacity of physical cameras. To overcome this, we propose a test-time spatial scaffolding framework that enhances policies without fine-tuning or hardware modifications, introducing a novel training-free paradigm. By synchronizing online simulation scenes with an interaction-aware state disambiguation mechanism, the method dynamically renders complementary virtual viewpoints to reveal critical spatial relationships and collaborates with frozen multimodal large models for control execution. Experimental results demonstrate that this framework significantly improves manipulation success rates in real-world scenarios. Notably, the success rate for plug insertion increases from 26.7% to 66.7%, while the Tower of Hanoi task achieves a breakthrough from 0% to 100%.
📝 Abstract
Frontier multimodal foundation models (e.g., GPT-6 Astra) have recently shown strong potential for direct robotic control, yet their performance on fine manipulation remains limited. We argue that an important source of failure is not necessarily insufficient policy capability, but insufficient spatial observability, where task-critical spatial relationships may be poorly revealed by the existing physical camera setup. We introduce SpatialHarness, a test-time embodied harness that provides test-time spatial scaffolding for fine robotic manipulation without policy fine-tuning or changes to the physical sensing setup. SpatialHarness maintains an online simulated scene synchronized with real-world execution, identifies task-critical spatial relationships, and renders complementary virtual views that expose them to a frozen multimodal policy. To keep the simulated scene aligned during interaction, we develop interaction-aware scene synchronization that distinguishes static, held, and transition modes. We evaluate SpatialHarness on four real-robot manipulation tasks spanning precise geometric alignment, object-relative placement, and articulated-object interaction. Using the same frozen GPT-6 Astra policy, SpatialHarness substantially improves task success, including from 26.7% to 66.7% on plug insertion and from 0% to 100% on Tower of Hanoi. These results indicate that improving spatial observability at test time can unlock fine-manipulation capabilities already present in strong multimodal foundation models. Project website: https://emilia113.github.io/SpatialHarness/.