🤖 AI Summary
This study addresses catastrophic forgetting in robotic continual learning caused by high storage overhead and unreliable trajectories generated by world action models. We propose a generative experience replay framework integrating reliability filtering with drift-aware replay. Specifically, this work pioneers the use of a frozen inverse dynamics model to evaluate action-visual consistency for selecting reliable trajectories, and adaptively samples high-drift historical data for replay training based on normalized predictive drift. Experiments on the LIBERO simulation benchmark demonstrate that our method surpasses existing state-of-the-art approaches, achieving an AUC of 90.97 on Goal tasks while retaining only 4.9% of historical steps, thereby significantly reducing storage costs.
📝 Abstract
Genuine embodied agency requires robots to turn continuous real-world experience into lasting, transferable skills. This demands continual learning that integrates new capabilities without eroding prior knowledge as tasks and environments evolve. Experience replay mitigates forgetting, but storing complete demonstrations becomes costly as tasks accumulate. World-action models offer a generative alternative, reconstructing past experience through joint predictions of actions and future observations. However, visually coherent rollouts may contain actions that cannot realize the predicted transitions, while new-task adaptation can disrupt previously learned behavior. RIFAR therefore combines reliability screening with drift-aware replay selection. It reconstructs trajectories from compact demonstration prefixes and uses a frozen inverse-dynamics model to assess action-visual consistency. Training first combines current demonstrations with the highest-quality screened trajectories. RIFAR then compares action predictions before and after this adaptation on identical historical inputs, reselecting trajectories with larger normalized drift from the same screened pool for continued training. Across three LIBERO suites and real-world experiments, RIFAR surpasses the previous state of the art in WAM-based generative replay. On LIBERO-Goal, it achieves 90.97 AUC while retaining only 320 historical time steps per task, approximately 4.9% of the steps retained using 50-demonstration replay.