🤖 AI Summary
This study addresses the challenges of sparse rewards and distribution mismatch arising from reliance on expert demonstrations in reinforcement learning for generalist robot policies. It proposes a demonstration-free reward learning paradigm, theoretically establishing that intermediate-step success probabilities can be recursively derived from terminal outcomes. By introducing the eVTA$_0$ model with a temporal difference bootstrapping mechanism, the authors construct RLER, a closed-loop framework that directly learns dense reward feedback from mixed-quality experience. This approach enables adaptive reward evolution without requiring expert demonstrations. Evaluated on the LIBERO benchmark, the method achieves average success rate improvements of 5.4%–13.8%, alongside gains of 20%–26% in real-world manipulation tasks and 35%–36% in out-of-distribution scenarios, demonstrating substantial robustness and generalization capabilities.
📝 Abstract
Reinforcement learning (RL) enables generalist robot policies to improve through trial-and-error interaction, yet its effectiveness is fundamentally constrained by sparse task rewards. Existing general-purpose reward models typically alleviate this issue by learning task progress from expert demonstrations, but introduce a distribution mismatch with the mixed-quality rollouts encountered during policy optimization, making their estimates unreliable on suboptimal and failed behaviors from which the policy must learn. In this work, we introduce a demonstration-free reward learning paradigm where dense reward feedback can be learned directly from sparse task outcomes and policy experience. We theoretically show that terminal task outcomes implicitly define dense success-probability feedback at intermediate timesteps, which can be recursively learned through bootstrapping. Based on this insight, we introduce eVTA$_0$, which learns success probabilities from mixed-quality policy rollouts through temporal-difference-style bootstrapping, without expert demonstrations or intermediate annotations. We further introduce RL with Evolving Rewards (RLER), a closed-loop framework that adapts eVTA$_0$ using newly collected rollouts as the policy evolves. Experiments show that eVTA$_0$ provides more informative rewards than state-of-the-art reward models and achieves the best average policy performance across all LIBERO task suites under the same RL training budget, improving success rates by 5.4%-13.8% over the initial policy. In real-world manipulation, RLER further improves overall success rates by 20%-26%, with 35%-36% gains under out-of-distribution conditions. These results demonstrate the effectiveness of demonstration-free reward learning and adapting rewards as the policy evolves. Project webpage: https://duowuyms.github.io/evta0.