HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出HIL-UMI框架,通过手持UMI演示和策略指导,在不依赖物理机器人的情况下解决VLA模型部署中的适应性问题。
📝 Abstract
Large-scale vision-language-action (VLA) models provide powerful priors for robot manipulation, yet adapting them to a specific deployment remains challenging. Supervised fine-tuning (SFT) on task-specific demonstrations provides a step toward deployment, but faces two persistent limitations: static data provide limited coverage of out-of-distribution states, and standard imitation objectives do not distinguish progressing behavior from less useful data. Interactive post-training can address these limitations, but typically requires repeated policy execution and human intervention on a physical robot. We introduce HIL-UMI, a policy-guided Universal Manipulation Interface (UMI) framework for robot-free human-in-the-loop VLA post-training. During handheld UMI demonstrations, HIL-UMI queries the current policy on the same observation stream without executing its predictions. The Energy Score compares the human action trajectory with policy inference and triggers collection when their discrepancy indicates an out-of-distribution region. In a separate feedback loop, low online advantage predictions identify essential segments for refining a progress-based advantage estimator. The updated estimator then guides advantage-conditioned behavioral cloning using a balanced mixture of base demonstrations and new policy data. This design preserves the iterative and policy-aware nature of human-in-the-loop learning while decoupling data collection from robot deployment. Experiments on four real-world tasks spanning long-horizon and precise manipulation show that HIL-UMI achieves consistent improvement over SFT and benefits from both targeted collection and advantage refinement. Moreover, HIL-UMI outperforms HG-DAgger on Clean Up Table with lower per-frame collection time, suggesting a scalable path for VLA post-training across operators and locations.
Problem

Research questions and friction points this paper is trying to address.

vision-language-action models
robot manipulation
supervised fine-tuning
out-of-distribution states
imitation objectives
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human-in-the-loop
Universal Manipulation Interface
Energy Score
Advantage Estimator
Policy-guided
Z
Zimu Han
Center on Frontier Computing Studies, School of Computer Science, Peking University, China; National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University, China; PrimeBot, China; Xi’an Jiaotong University, China
Y
Yiming Zeng
Center on Frontier Computing Studies, School of Computer Science, Peking University, China; National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University, China; PrimeBot, China; Xi’an Jiaotong University, China
Jiyao Zhang
Jiyao Zhang
Peking University
Embodied AIRobotics3D Vision
Z
Zihao Zhao
Center on Frontier Computing Studies, School of Computer Science, Peking University, China
Yuanfei Wang
Yuanfei Wang
Peking University
robot learningreinforcement learning
Yixiang Jin
Yixiang Jin
Samsung R&D Institute China - Beijing
RoboticsRobot LearningRobot Simulator
Shiqi Li
Shiqi Li
Texas A&M University
computer vision
S
Shuangben Chen
Center on Frontier Computing Studies, School of Computer Science, Peking University, China
Wei Huang
Wei Huang
Center on Frontier Computing Studies, School of Computer Science, Peking University, China
R
Ruodai Li
JD Technology, China
H
Hui Shen
JD Technology, China
H
Hao Dong
Center on Frontier Computing Studies, School of Computer Science, Peking University, China; National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University, China; PrimeBot, China