VLaRL: Augmenting Vision-Language-Action Models with Simulation-Trained Latent-Conditioned Residual RL

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited execution precision of Vision-Language-Action (VLA) models in contact-rich tasks, alongside the high cost and safety risks of real-world robotic reinforcement learning. We propose VLaRL, a framework that freezes a pretrained VLA and leverages its internal latent representations as control conditions and a sim-to-real transfer interface, thereby circumventing pixel-level alignment. By training a residual policy in simulation and employing a lightweight mapper to align cross-domain latent feature distributions, VLaRL achieves efficient transfer and zero-shot online deployment without adaptation. Experiments across four contact-rich tasks and two VLA backbones demonstrate that VLaRL significantly improves real-world success rates, validating the effectiveness of the latent conditioning and feature alignment mechanisms.
📝 Abstract
Vision-language-action (VLA) models provide broad, instruction-conditioned manipulation behaviors, but their physical execution can remain imprecise during contact-rich interaction. Residual reinforcement learning (RL) can correct such errors while keeping the VLA frozen, but real-robot RL is costly and safety-critical. We propose VLA Latent-Conditioned RL (VLaRL), which enables residual RL for frozen VLAs to be trained in simulation and deployed on real robots without real-world RL or online adaptation. The key challenge is transferring the learned residual policy despite the visual gap between simulation and reality. Rather than requiring pixel-level visual correspondence, VLaRL uses the VLA's internal vision-language latent representation to condition residual control and as the sim-to-real transfer interface, and learns a lightweight mapper that transforms simulation-derived latents toward the real latent distribution. Across four contact-rich manipulation tasks and two VLA backbones, VLaRL improves real-world success in all task-backbone combinations, while controlled ablations demonstrate the importance of both latent conditioning and latent alignment for transferring simulation-trained residual control.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
Residual reinforcement learning
Sim-to-real transfer
Contact-rich manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Residual Reinforcement Learning
Sim-to-Real Transfer
Latent Conditioning
Latent Alignment
🔎 Similar Papers
No similar papers found.