BEE: Intervention-Adaptive Real-World Reinforcement Learning with Vision-Language-Action Models

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究提出BEE框架,通过自适应人类干预优化真实世界中视觉-语言-动作模型的强化学习,提高任务成功率并减少所需的人类干预。
📝 Abstract
Vision-language-action (VLA) models handle long-horizon manipulation, yet success hinges on a few precision-critical phases where millimeter-scale errors undo all prior progress. Online reinforcement learning (RL) can optimize exactly these actions, but free exploration is far too costly on real robots, which makes human corrections indispensable. However, existing online RL methods for VLAs either cannot incorporate such corrections or fold them into undifferentiated supervision. Yet human corrections are not uniformly noisy but reliable along some action dimensions and variable along others. Building on this, we introduce BEE, an intervention-adaptive framework for real-world RL on a frozen VLA that lets the policy go BEyond Expert imitation. We formulate human corrections not as actions to reproduce but as evidence about a constraint: a Correction Model predicts how a human would correct a given VLA proposal and how consistent the correction is along each action dimension. This predicted consistency sets the per-dimension tightness of a constraint on policy optimization. Where corrections are consistent the policy stays close to the human, and where they vary, the constraint relaxes. We evaluate BEE on three real-world manipulation tasks and one LIBERO-Pro simulation task at a matched online-data budget. BEE attains the highest success rate on every task, 91.2% on average against 57.5% for RLT and 42.1% for DSRL, and the lowest human intervention rate on all real-world tasks.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
Reinforcement Learning
Human Corrections
Real-World Manipulation
Precision-Critical Phases
Innovation

Methods, ideas, or system contributions that make the work stand out.

intervention-adaptive
visual-language-action models
reinforcement learning
correction model
constraint-based optimization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Weihui Zhao
South China University of Technology
X
Xiaohan Yan
AgiBot
Z
Zunian Wan
AgiBot
X
Xuan Du
AgiBot
Z
Zhaozhan Chi
AgiBot
J
Jianbo Mao
AgiBot
R
Ruipu Wu
AgiBot
Rushuai Yang
Rushuai Yang
Hong Kong University of Science and Technology
Reinforcement LearningEmbodied AI
H
Houlin Li
AgiBot
S
Shukai Yang
AgiBot
J
Jing Wu
AgiBot
Y
Yuxiang Yan
AgiBot
Y
Yongcheng Liu
AgiBot
C
Chuankang Li
AgiBot
G
Guanghui Ren
AgiBot
W
Wei Shan
AgiBot
Maoqing Yao
Maoqing Yao
Google