ForceRFT: Refining VLA Actions through Force-Guided Residual Reinforcement Learning

📅 2026-09-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文提出ForceRFT方法,通过力引导的残差强化学习改进VLA策略,结合人类监督和自主任务结果优化接触依赖性修正,提高机器人任务成功率。
📝 Abstract
Force-conditioned vision-language-action (VLA) policies can respond to contact, but when trained solely on demonstrations, their recovery behavior may be limited by demonstration coverage, and they do not learn from deployment outcomes. Human corrective imitation provides additional recovery examples, but its objective matches local action targets without explicitly optimizing task return. We present ForceRFT, a force-guided residual reinforcement learning framework that learns contact-dependent corrections from human supervision and autonomous task outcomes. A frozen, demonstration-trained SmolVLA-based prior generates force-conditioned action chunks, while a lightweight residual actor refines individual end-effector pose commands using wrist feedback acquired during chunk execution. The decision-time wrist wrench, its temporal change, and the selected base motion condition both residual correction and value estimation. Human corrections supervise the residual actor, while verified autonomous transitions train the twin critics and support value-guided updates to the same actor. Bootstrapping is restricted to autonomous segments, preventing TD credit from crossing human-intervention boundaries. Real-robot experiments on plug insertion, ring-on-peg assembly, and whiteboard wiping show higher autonomous success rates than the evaluated demonstration-trained and residual-imitation baselines. Comparisons with residual imitation support value-guided residual optimization, while plug-insertion ablations indicate the benefit of direct execution-time wrist feedback.
Problem

Research questions and friction points this paper is trying to address.

Force-conditioned VLA
demonstration coverage
recovery behavior
human corrective imitation
task return
Innovation

Methods, ideas, or system contributions that make the work stand out.

force-guided
residual reinforcement learning
contact-dependent corrections
wrist feedback
value-guided
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yichen Wang
Yichen Wang
Huazhong University of Science and Technology
C
Chaoyang Zhang
Jiangsu Key Laboratory of Advanced Food Manufacturing Equipment and Technology, Jiangnan University, Wuxi, Jiangsu, China.
X
Xuqi Su
College of Information Science and Technology, Eastern Institute of Technology, Ningbo, Ningbo 315200, China.
J
Jun Ma
Robotics and Autonomous Systems Thrust, The Hong Kong University of Science and Technology (Guangzhou), Guangzhou 511453, China.
H
Haiyue Zhu
Department of Electrical and Computer Engineering, National University of Singapore, Singapore 117583.
X
Xiaocong Li
College of Information Science and Technology, Eastern Institute of Technology, Ningbo, Ningbo 315200, China.