decision-transformer offline rl

Designs and trains transformer-based sequence models that implement offline reinforcement learning by predicting actions conditioned on past state-action sequences and target returns using pre-collected demonstration or logged data. Builds the data preprocessing and training pipelines for Decision Transformer architectures and evaluates their data efficiency, reward-conditioning behavior, and fidelity of generated action sequences to desired behaviors.

decision-transformerofflinerl

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Offline Trajectory Generalization for Offline Reinforcement Learning

Apr 16, 2024
ZZ
Ziqi Zhao
🏛️ Shandong University | Leiden University | Westlake University

Offline reinforcement learning faces dual challenges: weak policy generalization and low-quality simulated trajectories. Existing model-based data augmentation methods are constrained by short-horizon simulation and lack mechanisms for evaluating or correcting generated data. This paper proposes OTTO, the first framework to introduce the Causal World Transformer into offline trajectory generalization. OTTO jointly models states and rewards, and designs four high-reward-oriented trajectory perturbation strategies to generate high-fidelity synthetic data. Its plug-and-play architecture integrates seamlessly with any underlying RL algorithm, enabling mixed training on original offline data and simulated trajectories without architectural modifications. Evaluated on the D4RL benchmark, OTTO achieves an average performance improvement of 12.7% over state-of-the-art methods. Notably, it demonstrates substantial gains in sparse-reward settings and out-of-distribution state generalization tasks.

Enhancing offline RL with long-horizon trajectory simulationEvaluating and correcting low-quality augmented dataImproving model-free offline RL performance in sparse-reward environments

Prompt-DT suffers from weak few-shot prompting capability and poor task discrimination in offline reinforcement learning, primarily due to data scarcity, high acquisition cost, or safety constraints. To address these challenges, this work introduces pre-trained language models (LLMs) into the Decision Transformer framework for the first time, proposing a tripartite synergistic mechanism: LLM-based initialization, LoRA-based fine-tuning, and prompt regularization. This design significantly enhances the model’s task awareness and generalization ability under limited prompt supervision. On the MuJoCo benchmark, our method achieves comparable performance to Prompt-DT trained on full prompt data while using only 10% of the prompts. Ablation studies confirm the necessity and effectiveness of each component. Overall, this work establishes a novel paradigm for prompt-based offline RL in low-data, high-safety-critical settings.

Addresses data scarcity and safety issues in offline reinforcement learning.Enhances few-shot prompt ability for unseen tasks in Decision Transformer.Improves task differentiation through language model initialization and regularization.

This work addresses a key limitation of conventional autoregressive approaches in offline contextual reinforcement learning, which merely imitate suboptimal behavioral policies and struggle to recover optimal ones. To overcome this, the authors propose the Decision Importance Transformer (DIT) framework, which introduces the actor-critic paradigm into contextual reinforcement learning for the first time. DIT leverages a Transformer architecture to model the advantage function over suboptimal trajectories and employs an advantage-weighted maximum likelihood objective to train the policy model. This approach effectively corrects policy bias present in historical data, thereby surpassing the performance ceiling of standard imitation learning. Empirical results demonstrate that DIT significantly outperforms existing baselines on both multi-armed bandit and Markov decision process tasks, exhibiting particularly strong policy optimization capabilities when trained on suboptimal datasets.

autoregressive transformerimitation learningin-context reinforcement learning

Existing offline safe reinforcement learning methods struggle with complex, multi-threaded, and temporally dependent real-world constraints. This paper proposes STL-Decision Transformer, the first approach to explicitly incorporate Signal Temporal Logic (STL) specifications as conditional inputs into the Decision Transformer architecture, enabling joint optimization of reward maximization and satisfaction of multi-granularity temporal safety constraints. By directly encoding STL semantics into the policy conditioning mechanism, our method overcomes fundamental limitations of conventional conditional policies in expressivity and generalizability for temporal logic, while supporting continuous, controllable adjustment of STL satisfaction degrees. Evaluated on the DSRL benchmark, STL-Decision Transformer significantly outperforms state-of-the-art methods, achieving simultaneous improvements in both cumulative reward and constraint satisfaction rate. These results demonstrate its effectiveness and robustness for offline, constraint-driven policy learning under rich temporal safety requirements.

Complex Multi-threaded ProblemsOffline Reinforcement LearningSequential Dependence

TOP-ERL: Transformer-based Off-Policy Episodic Reinforcement Learning

Oct 12, 2024
GL
Ge Li
🏛️ Karlsruhe Institute of Technology

To address the challenges of off-policy training difficulty, low sample efficiency, and unstable policy updates in episodic reinforcement learning (ERL), this paper proposes the first ERL framework supporting trajectory-level off-policy updates. Methodologically, it introduces three key innovations: (1) the first integration of Transformers into the critic architecture for ERL, enabling segmented modeling of long-horizon action trajectories; (2) a hybrid value estimation scheme combining *n*-step returns with segment-wise sequence value prediction, thereby relaxing strict on-policy constraints; and (3) policy parameterization via movement primitives to enhance interpretability and generalization. Evaluated on complex robotic control benchmarks, the framework achieves significant improvements over state-of-the-art methods. Ablation studies quantitatively validate the individual contributions of each component to sample efficiency, convergence stability, and final performance.

Complex TasksLearning EfficiencyReinforcement Learning

Latest Papers

What's happening recently
View more

A Comparison Between Decision Transformers and Traditional Offline Reinforcement Learning Algorithms

Nov 20, 2025
AM
Ali Murtaza Caunhye
🏛️ University of KwaZulu-Natal | Centre for Artificial Intelligence Research

This study investigates how decision transformers (DTs) compare to conventional offline reinforcement learning (RL) algorithms—specifically conservative Q-learning (CQL) and implicit Q-learning (IQL)—under varying reward densities (dense vs. sparse) in the ANT continuous-control benchmark. Method: We conduct a systematic, controlled empirical evaluation across uniformly configured offline datasets of varying quality and reward sparsity. Contribution/Results: We find that DTs exhibit remarkable robustness to reward density shifts: they outperform both CQL and IQL in sparse-reward regimes and under medium-quality offline data, achieving higher policy performance, greater stability, and lower evaluation variance. In contrast, IQL excels in dense-reward settings, while CQL demonstrates superior overall robustness across diverse conditions. Crucially, this work provides the first empirical evidence that sequence modeling—via autoregressive action prediction—confers distinct advantages in low signal-to-noise-ratio feedback environments. These findings offer principled guidance for reward-structure-aware algorithm selection and design in offline RL.

Analyzes algorithm sensitivity to reward density and data qualityCompares Decision Transformers with traditional offline RL algorithmsEvaluates performance in dense versus sparse reward environments

A New Perspective on Transformers in Online Reinforcement Learning for Continuous Control

Oct 15, 2025
NK
Nikita Kachaev
🏛️ Cognitive AI Lab | IAI MIPT

Prior work has identified instability in Transformer-based online model-free reinforcement learning (RL), primarily due to high sensitivity to policy/value network architecture, parameter sharing schemes, and temporal modeling strategies. Method: This paper presents the first systematic study of Transformer design for online continuous control, proposing a stable and efficient Actor-Critic architecture featuring serialized state inputs, temporal slicing, cross-network parameter sharing, and conditional input conditioning—unified to support both vector and image observations. Contribution/Results: The proposed method significantly improves training stability and generalization across diverse online RL benchmarks. It achieves state-of-the-art performance on both fully observed (e.g., MuJoCo) and partially observed (e.g., DeepMind Control Suite with proprioceptive+visual inputs) tasks. By providing a reproducible architectural blueprint and empirically validated design principles, this work establishes a new paradigm and practical guidelines for deploying Transformers in online RL settings.

Addressing architectural design challenges in actor-critic transformer networksDeveloping stable training strategies for transformers in continuous control tasksExploring transformer applications in online model-free reinforcement learning for control

This work addresses the input redundancy inherent in the Decision Transformer (DT) when utilizing Return-to-Go (RTG) sequences, which compromises both computational efficiency and performance. The authors propose the Decoupled Decision Transformer (DDT), which is the first to explicitly identify the redundancy in RTG sequences and decouple the RTG conditioning mechanism: only the most recent RTG value is used to guide action prediction, while the Transformer backbone processes solely the observation and action sequences. This streamlined architecture reduces unnecessary computation, enhances inference efficiency, and achieves significant performance improvements over the original DT across multiple offline reinforcement learning benchmarks, matching or surpassing the performance of current state-of-the-art DT variants.

Decision Transformeroffline reinforcement learningredundancy

This work addresses the inefficiency of the standard Decision Transformer, which embeds Return-to-Go (RTG) as standalone tokens in the autoregressive sequence, resulting in unnecessarily long sequences and high computational overhead. The authors propose decoupling the RTG conditioning from the autoregressive sequence and instead injecting it directly into the state representations. This enables the model to operate solely on compact (state, action) sequences, achieving a novel separation between sparse RTG signals and dense trajectory information. Evaluated on the D4RL benchmark, the proposed method significantly outperforms the standard Decision Transformer, attaining state-of-the-art performance while reducing sequence length by approximately one-third and substantially improving inference efficiency.

computational efficiencyDecision Transformeroffline reinforcement learning

This work addresses the challenge of achieving unified and scalable multitask decision-making across a vast array of heterogeneous reinforcement learning environments. It introduces LDM-v0, a Transformer-based universal policy model trained offline at scale on multimodal trajectory data from approximately 1,000 diverse domains—including robotics, autonomous driving, inventory management, cybersecurity, trading, and gaming. The model performs supervised next-action prediction conditioned on historical observations, actions, rewards, and termination signals. For the first time, it demonstrates that a single Transformer policy can match the performance of task-specific policies across over a thousand heterogeneous tasks, establishing a unified paradigm for large-scale multitask reinforcement learning.

heterogeneous environmentslarge decision modelsmulti-task reinforcement learning

Hot Scholars

ZC

Zhirong Chen

Master, Institute of Computing Technology, Chinese Academy of Sciences
Computer ArchitectureMachine Learning
BL

Bei Liu

Postdoc at HKUST
Speech ProcessingLarge Language ModelsEfficient AIModel Compression
LD

Li Dong

Microsoft Research
Natural Language Processing
DL

Dongsu Lee

University of Texas at Austin
Reinforcement Learning