Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of Vision-Language-Action (VLA) models in autonomously improving from failures and the reliance of real-world reinforcement learning on manual resets and supervision. To this end, we propose the FIND framework, which reformulates practice as a weakness-aware task selection problem conditioned on scene states and performance. This enables agents to autonomously identify weaknesses, select tasks, and leverage post-execution scene evaluations. By integrating a frozen π0.5 VLA model with residual off-policy reinforcement learning, the framework achieves continuous policy optimization without human-provided rewards. Experiments demonstrate that the system completes 456 autonomous trials within six hours—requiring minimal scene-recovery interventions—and improves success rates across eight tasks from 55% to 71.9%, effectively validating the feasibility of autonomous robotic evolution.
📝 Abstract
Vision--language--action (VLA) models provide strong priors for robotic manipulation but are typically deployed as frozen policies, unable to improve from their own failures. Real-world reinforcement learning (RL) offers a path to continued improvement, yet manual environment resets and task-success supervision hinder autonomous learning. We introduce \textbf{FIND}, an agentic real-world RL framework that closes the loop between scene understanding, weakness-aware practice, self-evaluation, and policy improvement in a persistent workspace. FIND reframes autonomous practice as a scene-conditioned, performance-aware task-selection problem: instead of restoring a predefined scene after each rollout, it uses the resulting scene to determine what to practice next. A vision--language agent identifies feasible tasks from a predefined library, prioritizes those with lower recent success rates, and evaluates outcomes using paired pre- and post-execution observations. We instantiate FIND with a frozen $\pi_{0.5}$ VLA and residual off-policy RL. Across eight real-world manipulation tasks, the independent human-assessed success rate improves from $55\%$ to $71.9\%$. A representative run completes 456 autonomous episodes within 6 hours of interaction, requiring 30 scene-recovery interventions and no human-provided reward labels during online learning. Ablations and systematic evaluations further examine key design choices, agent evaluation accuracy, and human intervention requirements. Our website is made publicly available at: FIND.github.io.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action Models
Real-World Reinforcement Learning
Autonomous Self-Improvement
Robotic Manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Real-World Reinforcement Learning
Autonomous Self-Improvement
Residual Off-Policy RL
Agentic Task Selection