SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出SynthDemo-RL框架,通过合成示例和强化学习解决VLA模型在稀疏奖励下难以探索成功轨迹的问题。
📝 Abstract
Fine-tuning Vision-Language-Action (VLA) models commonly relies on human teleoperation demonstrations, while reinforcement learning (RL) with sparse binary rewards faces an exploration challenge when successful trajectories are rarely sampled. We propose SynthDemo-RL, a teacher-student framework in which an automated teacher converts simulator-privileged state into successful manipulation trajectories, a VLA student is distilled from them by supervised fine-tuning (SFT), and PPO with binary task-success rewards refines the student. We study reward coverage, the fraction of tasks for which at least one success is observed under the fixed evaluation protocol, as a complement to the average success rate. On LIBERO-PRO, a public benchmark of perturbed LIBERO tasks for which no demonstrations exist, 27 of 57 scored tasks are at exactly 0% success for a pi_0.5 policy fine-tuned on the original LIBERO tasks. Direct PPO from this policy, under the same PPO recipe and the same RL compute as SynthDemo-RL's refinement stage, rescues 10 of these 27 tasks and leaves 17 at 0%. SynthDemo-RL, with 50 synthesized trajectories per task and no new human demonstrations, rescues all 27 and reaches average success rates of 97.8% and 97.1% on the Position and Task axes of LIBERO-PRO, respectively. On standard LIBERO, the same pipeline reaches 96.0% with no human demonstrations, within 1.7 points of pi_0.5 trained on 50 human demonstrations per task. We further validate the pipeline on RoboTwin 2.0 and verify that trajectories from a policy trained in a MuJoCo twin execute open-loop on a physical robot.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action (VLA)
reinforcement learning (RL)
sparse binary rewards
exploration challenge
successful trajectories
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synthetic Demonstrations
Reinforcement Learning
Vision-Language-Action Models
Automatic Teacher
Policy Optimization
🔎 Similar Papers
No similar papers found.
H
Hiroaki Kingetsu
Fujitsu Limited, Kawasaki, Japan.
H
Hiroaki Kurihara
Fujitsu Limited, Kawasaki, Japan.
K
Kaoru Yokoo
Fujitsu Limited, Kawasaki, Japan.
Kenji Fukumizu
Kenji Fukumizu
The Institute of Statistical Mathematics
Machine learningstatistics
Manohar Kaul
Manohar Kaul
Fujitsu Limited, Kawasaki, Japan.