Predictive Coding for Decision Transformer

📅 2024-10-04
🏛️ IEEE/RJS International Conference on Intelligent RObots and Systems
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Decision Transformers (DTs) underperform in sparse-reward, long-horizon offline goal-directed reinforcement learning, primarily because scalar return conditioning fails to capture temporal compositionality. To address this, we propose Predictive Coding Conditioning (PCC), a novel conditioning mechanism that replaces scalar returns with generalized future-state or goal representations—marking the first integration of predictive coding into the DT framework. PCC enables joint policy modeling conditioned on both past observations and future goals. Our method extends the Transformer architecture with predictive coding–based representations, future-conditioned modeling, and offline behavior cloning training. Evaluated on eight challenging offline benchmarks across AntMaze and FrankaKitchen, PCC matches or surpasses state-of-the-art value-based and Transformer-based baselines. Furthermore, it demonstrates strong generalization in real-world physical robot goal-reaching tasks.

Technology Category

Cognitive Modeling & Cognitive Systems: Neural Spike CodingHumans and AI: Human-Aware Planning and Behavior PredictionReasoning under Uncertainty: Relational Probabilistic Models

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendation
📝 Abstract
Recent work in offline reinforcement learning (RL) has demonstrated the effectiveness of formulating decision-making as return-conditioned supervised learning. Notably, the decision transformer (DT) architecture has shown promise across various domains. However, despite its initial success, DTs have underperformed on several challenging datasets in goal-conditioned RL. This limitation stems from the inefficiency of return conditioning for guiding policy learning, particularly in unstructured and suboptimal datasets, resulting in DTs failing to effectively learn temporal compositionality. Moreover, this problem might be further exacerbated in long-horizon sparse-reward tasks. To address this challenge, we propose the Predictive Coding for Decision Transformer (PCDT) framework, which leverages generalized future conditioning to enhance DT methods. PCDT utilizes an architecture that extends the DT framework, conditioned on predictive codings, enabling decision-making based on both past and future factors, thereby improving generalization. Through extensive experiments on eight datasets from the AntMaze and FrankaKitchen environments, our proposed method achieves performance on par with or surpassing existing popular value-based and transformer-based methods in offline goal-conditioned RL. Furthermore, we also evaluate our method on a goal-reaching task with a physical robot.
Problem

Research questions and friction points this paper is trying to address.

Improves decision transformers in offline RL tasks
Addresses inefficiency in return conditioning methods
Enhances performance in sparse-reward long-horizon tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Predictive coding enhances Decision Transformer
Generalized future conditioning improves policy learning
Combines past and future factors for decisions
🔎 Similar Papers
No similar papers found.
KAIST
T
T. Luu
School of Electrical Engineering, KAIST (Korea Advanced Institute of Science and Technology), Daejeon, Republic of Korea
D
Donghoon Lee
School of Electrical Engineering, KAIST (Korea Advanced Institute of Science and Technology), Daejeon, Republic of Korea
Chang D. Yoo
Chang D. Yoo
kaist
machine learningcomputer visionsignal processingspeech enhancementspeech recognition