🤖 AI Summary
This study addresses the heavy reliance of Vision-Language-Action (VLA) models on extensive demonstrations and the high cost of pure-RGB control by proposing the Robo-Harness K1 framework. Its core innovation lies in encapsulating 3D perception as a tool interface, enabling agents to select universal actions by querying geometric evidence such as depth and anchor points, thereby achieving robotic manipulation without modifying or fine-tuning the underlying VLM architecture. The proposed method attains an accuracy of 88.9% on the LIBERO-PRO benchmark, significantly outperforming pure-RGB baselines while demonstrating strong robustness against perturbations. Furthermore, through policy distillation, it surpasses OpenVLA under few-shot settings and achieves zero-fine-tuning cross-embodiment generalization across different robotic arms, validating the effectiveness of this perception-augmented paradigm.
📝 Abstract
Foundation vision-language models (VLMs) understand objects, instructions, and spatial relations, yet translating this capability into robotic manipulation remains difficult. Vision-language-action (VLA) models require extensive demonstrations and may compromise pretrained understanding, while direct RGB-only VLM control is costly and strongly dependent on model capability. We introduce Robo-Harness K1, a robot-use agent (RUA) framework that exposes perception as tools. The agent queries calibrated depth, persistent visual anchors, spatial measurements, and grasp hypotheses, then selects generic motions from the returned evidence. This interface makes 3D geometry accessible without changing the VLM architecture or training a depth encoder. On matched LIBERO-PRO tasks, Gemini 3.7 Flash with K1 reaches 77.8% accuracy, surpassing GPT-6 Astra's 61.1% with an RGB-only harness; K1 further improves Astra to 88.9%. Without target fine-tuning, Gemini with K1 transfers to three RoboSuite arms and dual-arm RoboTwin tasks. On RoboTwin, it achieves 32.0% on Easy and 28.0% on Hard, showing resilience to visual and environmental perturbations. K1 also produces tool-call traces aligned with next-token training. A Qwen3.5-9B student trained on only 107 teacher episodes reaches 44.2% accuracy on new initial states versus 30.2% for OpenVLA, and 13.9% on held-out task conditions versus 0.0% for OpenVLA. These results suggest that perception-augmented RUAs offer a promising route to sample-efficient, generalizable robotic policies that leverage VLM capabilities through an accessible tool interface.