🤖 AI Summary
This study addresses the coarse granularity of trajectory-level advantage estimation in embodied reinforcement learning, where failure penalties inadvertently penalize early effective actions. To this end, we propose a physics-relation-based cross-trajectory action chunk credit assignment method that, for the first time, infers action contributions solely from terminal outcomes. A confidence gating mechanism is further designed to mitigate noise interference. By integrating physics relation graphs with vision-language-action policy post-training, our approach refines terminal supervision signals without requiring auxiliary evaluators. Extensive experiments on the LIBERO benchmark and real-world robotic platforms demonstrate that the proposed method significantly improves task success rates and accelerates learning convergence.
📝 Abstract
Outcome-based reinforcement learning (RL) post-trains vision--language--action policies using terminal success signals, but assigns the same trajectory-level advantage to every action chunk. A failed episode can thus penalize useful early actions as if they caused the failure. Existing approaches seek finer-grained feedback through learned evaluators, adding task-specific supervision or additional model training. We explore, for the first time to our knowledge, whether physical relations across trajectories can provide action-chunk credit in embodied RL from terminal outcomes alone, without an auxiliary evaluator. The key insight is that rollouts reaching corresponding physical situations can serve as references for one another: their terminal outcomes provide evidence for assessing local progress. We introduce Physical Relations for Inferring Credit from Episodes(PRICE), with two components: (i) a physical relational graph that pools current and historical outcomes at corresponding chunk boundaries to estimate success potentials; and (ii) confidence-gated credit assignment that uses changes in these potentials to refine trajectory-level supervision. Our analysis connects oracle potential changes to the terminal-success objective and provides a finite-sample directional bound for outcome-independent evidence pools. Independent continuation tests show that PRICE's retained credits align with local progress, while experiments on LIBERO, RoboTwin 2.0, and real robots demonstrate improved task success over outcome-based baselines and faster learning.