Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of sparse rewards and distribution mismatch arising from reliance on expert demonstrations in reinforcement learning for generalist robot policies. It proposes a demonstration-free reward learning paradigm, theoretically establishing that intermediate-step success probabilities can be recursively derived from terminal outcomes. By introducing the eVTA$_0$ model with a temporal difference bootstrapping mechanism, the authors construct RLER, a closed-loop framework that directly learns dense reward feedback from mixed-quality experience. This approach enables adaptive reward evolution without requiring expert demonstrations. Evaluated on the LIBERO benchmark, the method achieves average success rate improvements of 5.4%–13.8%, alongside gains of 20%–26% in real-world manipulation tasks and 35%–36% in out-of-distribution scenarios, demonstrating substantial robustness and generalization capabilities.
📝 Abstract
Reinforcement learning (RL) enables generalist robot policies to improve through trial-and-error interaction, yet its effectiveness is fundamentally constrained by sparse task rewards. Existing general-purpose reward models typically alleviate this issue by learning task progress from expert demonstrations, but introduce a distribution mismatch with the mixed-quality rollouts encountered during policy optimization, making their estimates unreliable on suboptimal and failed behaviors from which the policy must learn. In this work, we introduce a demonstration-free reward learning paradigm where dense reward feedback can be learned directly from sparse task outcomes and policy experience. We theoretically show that terminal task outcomes implicitly define dense success-probability feedback at intermediate timesteps, which can be recursively learned through bootstrapping. Based on this insight, we introduce eVTA$_0$, which learns success probabilities from mixed-quality policy rollouts through temporal-difference-style bootstrapping, without expert demonstrations or intermediate annotations. We further introduce RL with Evolving Rewards (RLER), a closed-loop framework that adapts eVTA$_0$ using newly collected rollouts as the policy evolves. Experiments show that eVTA$_0$ provides more informative rewards than state-of-the-art reward models and achieves the best average policy performance across all LIBERO task suites under the same RL training budget, improving success rates by 5.4%-13.8% over the initial policy. In real-world manipulation, RLER further improves overall success rates by 20%-26%, with 35%-36% gains under out-of-distribution conditions. These results demonstrate the effectiveness of demonstration-free reward learning and adapting rewards as the policy evolves. Project webpage: https://duowuyms.github.io/evta0.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Generalist Robot Policies
Sparse Rewards
Reward Learning
Distribution Mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Demonstration-Free Reward Learning
Success-Probability Estimation
Temporal-Difference Bootstrapping
Evolving Rewards
Generalist Robot Policies
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
D
Duo Wu
Tsinghua University
H
Haifeng Wang
Tsinghua University, Yuanxing Robotics
Rongwei Lu
Rongwei Lu
Tsinghua University
Distributed machine learninggradient compressionfederated learning
J
Jinghe Wang
Tsinghua University
T
Tianyi Xiong
Tsinghua University
Z
Zhimin Wang
Tsinghua University
C
Chao Yu
Tsinghua University
S
Shuai Ma
Pengcheng Laboratory
Zhi Wang
Zhi Wang
Associate Professor, SIGS, Tsinghua University
multimedia networkedge computingdistributed machine learning