🤖 AI Summary
This study addresses the susceptibility of imitation learning to reinforcing errors and the unreliable value estimation in offline reinforcement learning when robots learn from mixed-quality experience. To overcome these challenges, this work proposes Predictive Action Chunking Learning (PACL). Its core innovation lies in designing a predictive action chunk evaluator that integrates future latent variable prediction with temporal difference augmentation, discretizing continuous Q-values into quality conditions to guide the joint training of diffusion policies beyond the limitations of conventional offline RL. Experimental results demonstrate that PACL enables continuous improvement without human intervention across both simulated and real-world robotic manipulation tasks, significantly outperforming mainstream baselines.
📝 Abstract
Robot policies deployed in real environments naturally accumulate mixed-quality experience, including successful executions, partial progress, and failures. Although these rollouts provide valuable information for further learning, directly incorporating them into imitation learning may reinforce undesirable behaviors, while offline reinforcement learning often suffers from unreliable value estimation under sparse rewards and limited data coverage. We consider a practical post-deployment setting where learning relies only on naturally accumulated autonomous rollouts, without additional human corrections or exploratory interaction. To effectively exploit such experience, we propose Predictive Action Chunk Learning (PACL). PACL first learns a predictive chunk-level critic that evaluates temporally extended action sequences and augments temporal difference learning with future latent prediction, providing richer supervision for long-horizon value estimation. The learned critic then converts chunk-level Q-values into discrete quality conditions, which guide a diffusion actor to learn jointly from these mixed-quality experiences without treating all behaviors as equivalent supervision. At inference, the actor generates multiple action chunks and the critic selects the highest valued candidate. Experiments across simulated and real-world robot manipulation tasks show that PACL consistently improves the pretrained policy and outperforms strong imitation learning and offline reinforcement learning baselines.