How (and How Not) to Use Data Augmentation in VLA Post-Training

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the generalization bottleneck in Vision-Language-Action (VLA) models caused by out-of-distribution visual shifts during reinforcement learning post-training, systematically investigating the underlying mechanisms of image augmentation. We find that indiscriminately augmenting the actor induces training collapse. Accordingly, we propose an asymmetric augmentation paradigm that applies data augmentation exclusively to the critic module while preserving clean inputs for the actor. Comparative experiments based on the PPO algorithm demonstrate that this strategy effectively enhances post-training robustness. On the LIBERO-Plus benchmark, our method improves out-of-distribution success rates by 7.8 and 10.0 percentage points for the π0.5 and GR00T N1.5 models, respectively, validating its effectiveness.
📝 Abstract
Vision-language-action (VLA) models currently demonstrate strong performance in a wide range of real-world robotics tasks. However, they often still lack the generalization ability to handle large visual out-of-distribution shifts. Post-training of VLAs with reinforcement learning (RL) has been shown to benefit robustness, but significant room for improvement remains. In this work, we systematically study the effect of image augmentation on VLA post-training. We find that it is crucial to augment only the critic module during RL updates, while leaving the actor's input clean during both rollouts and updates. For $\pi_{0.5}$ and GR00T N1.5 this raises out-of-distribution success on LIBERO-Plus by $7.8$ and $10.0$ points respectively, while augmenting the actor collapses training entirely. We investigate a range of augmentation types and strengths, and provide practical recommendations for improving generalization in VLA post-training.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
Data augmentation
Post-training
Out-of-distribution generalization
Reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action (VLA)
Data Augmentation
Reinforcement Learning
Post-training
Out-of-Distribution Generalization
🔎 Similar Papers
No similar papers found.