Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited out-of-distribution generalization of vision-language-action models following reinforcement learning fine-tuning. To overcome this limitation, we propose DRIVE, a novel framework that explicitly optimizes the behavioral diversity of successful trajectories for the first time. Specifically, DRIVE generates intrinsic rewards through temporally aligned trajectory comparisons, guiding policy exploration toward broader coverage of feasible solutions while preventing spurious reward signals arising from failures or superficial temporal discrepancies. Extensive evaluations across multiple simulation benchmarks and a real-world bimanual robotic platform demonstrate that our approach substantially enhances out-of-distribution generalization, yielding an average improvement in success rate of up to 9.2 percentage points over existing methods.
📝 Abstract
Reinforcement learning (RL) fine-tuning improves vision-language-action (VLA) policies through closed-loop experience, yet generalization beyond the fine-tuning distribution remains limited. Our analysis reveals a selective reshaping of exploration: RL contracts behavior globally, yet diversifies successful trajectories, elicits success with fewer rollouts, and covers more of the latent task-valid solution space than supervised fine-tuning. Broader successful-mode coverage may provide alternative strategies under distribution shifts. Inspired by this, we introduce DRIVE (Diversity-driven RL fIne-tuning for VLA gEneralization), which turns successful-behavior diversity into an explicit RL objective. DRIVE groups rollouts under matched task conditions, compares their trajectories with temporal alignment, and derives a success-conditioned intrinsic reward from relative behavioral diversity. This design encourages broader coverage of feasible solutions without rewarding diverse failures or superficial timing differences. Across LIBERO-Plus, ManiSkill3, and RoboTwin 2.0, DRIVE improves the average out-of-domain (OOD) performance over vanilla RL fine-tuning by 5.3 points on $π_0$ and 2.0 points on $π_{0.5}$. On a dual-arm AgileX PiPER-X platform, DRIVE further increases average OOD success from 64.1% to 73.3% (+9.2 points), demonstrating gains that persist under physical deployment.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action (VLA)
Reinforcement Learning Fine-Tuning
Generalization
Out-of-Distribution (OOD)
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning Fine-Tuning
Vision-Language-Action (VLA)
Behavioral Diversity
Intrinsic Reward
Out-of-Domain Generalization
🔎 Similar Papers
No similar papers found.