Robo-Harness K1: Harnessing Robot-Use Agents via Perception Augmentation

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the heavy reliance of Vision-Language-Action (VLA) models on extensive demonstrations and the high cost of pure-RGB control by proposing the Robo-Harness K1 framework. Its core innovation lies in encapsulating 3D perception as a tool interface, enabling agents to select universal actions by querying geometric evidence such as depth and anchor points, thereby achieving robotic manipulation without modifying or fine-tuning the underlying VLM architecture. The proposed method attains an accuracy of 88.9% on the LIBERO-PRO benchmark, significantly outperforming pure-RGB baselines while demonstrating strong robustness against perturbations. Furthermore, through policy distillation, it surpasses OpenVLA under few-shot settings and achieves zero-fine-tuning cross-embodiment generalization across different robotic arms, validating the effectiveness of this perception-augmented paradigm.
📝 Abstract
Foundation vision-language models (VLMs) understand objects, instructions, and spatial relations, yet translating this capability into robotic manipulation remains difficult. Vision-language-action (VLA) models require extensive demonstrations and may compromise pretrained understanding, while direct RGB-only VLM control is costly and strongly dependent on model capability. We introduce Robo-Harness K1, a robot-use agent (RUA) framework that exposes perception as tools. The agent queries calibrated depth, persistent visual anchors, spatial measurements, and grasp hypotheses, then selects generic motions from the returned evidence. This interface makes 3D geometry accessible without changing the VLM architecture or training a depth encoder. On matched LIBERO-PRO tasks, Gemini 3.7 Flash with K1 reaches 77.8% accuracy, surpassing GPT-6 Astra's 61.1% with an RGB-only harness; K1 further improves Astra to 88.9%. Without target fine-tuning, Gemini with K1 transfers to three RoboSuite arms and dual-arm RoboTwin tasks. On RoboTwin, it achieves 32.0% on Easy and 28.0% on Hard, showing resilience to visual and environmental perturbations. K1 also produces tool-call traces aligned with next-token training. A Qwen3.5-9B student trained on only 107 teacher episodes reaches 44.2% accuracy on new initial states versus 30.2% for OpenVLA, and 13.9% on held-out task conditions versus 0.0% for OpenVLA. These results suggest that perception-augmented RUAs offer a promising route to sample-efficient, generalizable robotic policies that leverage VLM capabilities through an accessible tool interface.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Robotic Manipulation
Vision-Language-Action Models
Robot-Use Agents
Perception Augmentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Perception Augmentation
Robot-Use Agent
Vision-Language Model
Tool Interface
Knowledge Distillation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Zexi Li
Zexi Li
Alibaba Group
Deep LearningLarge Language ModelsFederated Learning
Y
Yehang Zhang
The Hong Kong University of Science and Technology (Guangzhou)
W
Wenqian Li
The Chinese University of Hong Kong
H
Haojian Huang
The Hong Kong University of Science and Technology (Guangzhou)
C
Chenxu Wang
Knowin AI
S
Shiyuan Deng
Knowin AI
Y
Yangkai Wei
Knowin AI
T
Tianyi Zhang
Knowin AI
Binghui Xie
Binghui Xie
Knowin AI
B
Bohan Zhou
The Chinese University of Hong Kong
Y
Yifan Chang
Knowin AI
K
Kaiwen Zhou
Knowin AI
Ying-Cong Chen
Ying-Cong Chen
Hong Kong University of Science and Technology (Guangzhou)
Computer Vision and Pattern Recognition
J
James Cheng
The Chinese University of Hong Kong
Yinchuan Li
Yinchuan Li
Principal Researcher, Noah's Ark Lab
Generative ModelsEmbodied AIArtificial Intelligence