Rho: A Foundation for Efficiently Adaptable VLA Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that existing general-purpose physical AI models struggle to simultaneously achieve vision-language capabilities, precise bimanual control, and efficient adaptation to downstream tasks. We propose an open-source family of Vision-Language-Action (VLA) models that introduces a novel embodiment mid-training paradigm to bridge cross-platform discrepancies. Furthermore, we design a lightweight online latent policy learning algorithm based on corrective feedback, integrated with a Flow Matching action expert to enable precise control. The proposed approach outperforms existing open-source VLA models across three distinct bimanual platforms. Notably, it effectively handles edge cases with only 15 corrections, significantly reducing data requirements. All model weights have been publicly released.
📝 Abstract
General-purpose physical AI models must combine broad visual and linguistic capabilities with precise control across robot embodiments and efficient adaptation to downstream tasks. We introduce Rho, a family of open-weights VLA models for bimanual manipulation designed for data-light task adaptation on 3 embodiments representative of dual-arm robots across research labs and the industry -- YAM Box, UR AI Trainer, and FR3 Duo. We systematically ablate Rho's action-expert architecture and training recipe, and show in controlled simulation and physical-robot experiments that embodiment midtraining improves downstream adaptation. The resulting Rho variants for YAM Box, UR AI Trainer, and FR3 Duo match or outperform existing open-weights VLAs and achieve the strongest overall performance across the tasks, embodiments, and baselines evaluated in this report. We further demonstrate the Rho model family's built-in capacity for online adaptation: an internal latent policy learns from corrective feedback to select observation-conditioned noise inputs for the frozen flow-matching action expert. With as few as 15 corrected episodes, adapting this lightweight module enables Rho to handle task situations at the fringe of its offline finetuning distribution. Together, these results position Rho as both a strong general-purpose robotic manipulation model and a practical foundation for adaptation. We release the base Rho model and the embodiment-specific checkpoints to facilitate Rho's deployment in research experiments and practical industrial use cases.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
bimanual manipulation
task adaptation
robot embodiment
data-efficient learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action (VLA)
action-expert architecture
embodiment midtraining
online adaptation
flow-matching
🔎 Similar Papers
No similar papers found.