🤖 AI Summary
Decision Transformers (DTs) underperform in sparse-reward, long-horizon offline goal-directed reinforcement learning, primarily because scalar return conditioning fails to capture temporal compositionality. To address this, we propose Predictive Coding Conditioning (PCC), a novel conditioning mechanism that replaces scalar returns with generalized future-state or goal representations—marking the first integration of predictive coding into the DT framework. PCC enables joint policy modeling conditioned on both past observations and future goals. Our method extends the Transformer architecture with predictive coding–based representations, future-conditioned modeling, and offline behavior cloning training. Evaluated on eight challenging offline benchmarks across AntMaze and FrankaKitchen, PCC matches or surpasses state-of-the-art value-based and Transformer-based baselines. Furthermore, it demonstrates strong generalization in real-world physical robot goal-reaching tasks.
📝 Abstract
Recent work in offline reinforcement learning (RL) has demonstrated the effectiveness of formulating decision-making as return-conditioned supervised learning. Notably, the decision transformer (DT) architecture has shown promise across various domains. However, despite its initial success, DTs have underperformed on several challenging datasets in goal-conditioned RL. This limitation stems from the inefficiency of return conditioning for guiding policy learning, particularly in unstructured and suboptimal datasets, resulting in DTs failing to effectively learn temporal compositionality. Moreover, this problem might be further exacerbated in long-horizon sparse-reward tasks. To address this challenge, we propose the Predictive Coding for Decision Transformer (PCDT) framework, which leverages generalized future conditioning to enhance DT methods. PCDT utilizes an architecture that extends the DT framework, conditioned on predictive codings, enabling decision-making based on both past and future factors, thereby improving generalization. Through extensive experiments on eight datasets from the AntMaze and FrankaKitchen environments, our proposed method achieves performance on par with or surpassing existing popular value-based and transformer-based methods in offline goal-conditioned RL. Furthermore, we also evaluate our method on a goal-reaching task with a physical robot.