Score
Designs and trains transformer-based sequence models that implement offline reinforcement learning by predicting actions conditioned on past state-action sequences and target returns using pre-collected demonstration or logged data. Builds the data preprocessing and training pipelines for Decision Transformer architectures and evaluates their data efficiency, reward-conditioning behavior, and fidelity of generated action sequences to desired behaviors.
Offline reinforcement learning faces dual challenges: weak policy generalization and low-quality simulated trajectories. Existing model-based data augmentation methods are constrained by short-horizon simulation and lack mechanisms for evaluating or correcting generated data. This paper proposes OTTO, the first framework to introduce the Causal World Transformer into offline trajectory generalization. OTTO jointly models states and rewards, and designs four high-reward-oriented trajectory perturbation strategies to generate high-fidelity synthetic data. Its plug-and-play architecture integrates seamlessly with any underlying RL algorithm, enabling mixed training on original offline data and simulated trajectories without architectural modifications. Evaluated on the D4RL benchmark, OTTO achieves an average performance improvement of 12.7% over state-of-the-art methods. Notably, it demonstrates substantial gains in sparse-reward settings and out-of-distribution state generalization tasks.
Prompt-DT suffers from weak few-shot prompting capability and poor task discrimination in offline reinforcement learning, primarily due to data scarcity, high acquisition cost, or safety constraints. To address these challenges, this work introduces pre-trained language models (LLMs) into the Decision Transformer framework for the first time, proposing a tripartite synergistic mechanism: LLM-based initialization, LoRA-based fine-tuning, and prompt regularization. This design significantly enhances the model’s task awareness and generalization ability under limited prompt supervision. On the MuJoCo benchmark, our method achieves comparable performance to Prompt-DT trained on full prompt data while using only 10% of the prompts. Ablation studies confirm the necessity and effectiveness of each component. Overall, this work establishes a novel paradigm for prompt-based offline RL in low-data, high-safety-critical settings.
This work addresses a key limitation of conventional autoregressive approaches in offline contextual reinforcement learning, which merely imitate suboptimal behavioral policies and struggle to recover optimal ones. To overcome this, the authors propose the Decision Importance Transformer (DIT) framework, which introduces the actor-critic paradigm into contextual reinforcement learning for the first time. DIT leverages a Transformer architecture to model the advantage function over suboptimal trajectories and employs an advantage-weighted maximum likelihood objective to train the policy model. This approach effectively corrects policy bias present in historical data, thereby surpassing the performance ceiling of standard imitation learning. Empirical results demonstrate that DIT significantly outperforms existing baselines on both multi-armed bandit and Markov decision process tasks, exhibiting particularly strong policy optimization capabilities when trained on suboptimal datasets.
Existing offline safe reinforcement learning methods struggle with complex, multi-threaded, and temporally dependent real-world constraints. This paper proposes STL-Decision Transformer, the first approach to explicitly incorporate Signal Temporal Logic (STL) specifications as conditional inputs into the Decision Transformer architecture, enabling joint optimization of reward maximization and satisfaction of multi-granularity temporal safety constraints. By directly encoding STL semantics into the policy conditioning mechanism, our method overcomes fundamental limitations of conventional conditional policies in expressivity and generalizability for temporal logic, while supporting continuous, controllable adjustment of STL satisfaction degrees. Evaluated on the DSRL benchmark, STL-Decision Transformer significantly outperforms state-of-the-art methods, achieving simultaneous improvements in both cumulative reward and constraint satisfaction rate. These results demonstrate its effectiveness and robustness for offline, constraint-driven policy learning under rich temporal safety requirements.
To address the challenges of off-policy training difficulty, low sample efficiency, and unstable policy updates in episodic reinforcement learning (ERL), this paper proposes the first ERL framework supporting trajectory-level off-policy updates. Methodologically, it introduces three key innovations: (1) the first integration of Transformers into the critic architecture for ERL, enabling segmented modeling of long-horizon action trajectories; (2) a hybrid value estimation scheme combining *n*-step returns with segment-wise sequence value prediction, thereby relaxing strict on-policy constraints; and (3) policy parameterization via movement primitives to enhance interpretability and generalization. Evaluated on complex robotic control benchmarks, the framework achieves significant improvements over state-of-the-art methods. Ablation studies quantitatively validate the individual contributions of each component to sample efficiency, convergence stability, and final performance.
This study investigates how decision transformers (DTs) compare to conventional offline reinforcement learning (RL) algorithms—specifically conservative Q-learning (CQL) and implicit Q-learning (IQL)—under varying reward densities (dense vs. sparse) in the ANT continuous-control benchmark. Method: We conduct a systematic, controlled empirical evaluation across uniformly configured offline datasets of varying quality and reward sparsity. Contribution/Results: We find that DTs exhibit remarkable robustness to reward density shifts: they outperform both CQL and IQL in sparse-reward regimes and under medium-quality offline data, achieving higher policy performance, greater stability, and lower evaluation variance. In contrast, IQL excels in dense-reward settings, while CQL demonstrates superior overall robustness across diverse conditions. Crucially, this work provides the first empirical evidence that sequence modeling—via autoregressive action prediction—confers distinct advantages in low signal-to-noise-ratio feedback environments. These findings offer principled guidance for reward-structure-aware algorithm selection and design in offline RL.
Prior work has identified instability in Transformer-based online model-free reinforcement learning (RL), primarily due to high sensitivity to policy/value network architecture, parameter sharing schemes, and temporal modeling strategies. Method: This paper presents the first systematic study of Transformer design for online continuous control, proposing a stable and efficient Actor-Critic architecture featuring serialized state inputs, temporal slicing, cross-network parameter sharing, and conditional input conditioning—unified to support both vector and image observations. Contribution/Results: The proposed method significantly improves training stability and generalization across diverse online RL benchmarks. It achieves state-of-the-art performance on both fully observed (e.g., MuJoCo) and partially observed (e.g., DeepMind Control Suite with proprioceptive+visual inputs) tasks. By providing a reproducible architectural blueprint and empirically validated design principles, this work establishes a new paradigm and practical guidelines for deploying Transformers in online RL settings.
This work addresses the input redundancy inherent in the Decision Transformer (DT) when utilizing Return-to-Go (RTG) sequences, which compromises both computational efficiency and performance. The authors propose the Decoupled Decision Transformer (DDT), which is the first to explicitly identify the redundancy in RTG sequences and decouple the RTG conditioning mechanism: only the most recent RTG value is used to guide action prediction, while the Transformer backbone processes solely the observation and action sequences. This streamlined architecture reduces unnecessary computation, enhances inference efficiency, and achieves significant performance improvements over the original DT across multiple offline reinforcement learning benchmarks, matching or surpassing the performance of current state-of-the-art DT variants.
This work addresses the inefficiency of the standard Decision Transformer, which embeds Return-to-Go (RTG) as standalone tokens in the autoregressive sequence, resulting in unnecessarily long sequences and high computational overhead. The authors propose decoupling the RTG conditioning from the autoregressive sequence and instead injecting it directly into the state representations. This enables the model to operate solely on compact (state, action) sequences, achieving a novel separation between sparse RTG signals and dense trajectory information. Evaluated on the D4RL benchmark, the proposed method significantly outperforms the standard Decision Transformer, attaining state-of-the-art performance while reducing sequence length by approximately one-third and substantially improving inference efficiency.
This work addresses the challenge of achieving unified and scalable multitask decision-making across a vast array of heterogeneous reinforcement learning environments. It introduces LDM-v0, a Transformer-based universal policy model trained offline at scale on multimodal trajectory data from approximately 1,000 diverse domains—including robotics, autonomous driving, inventory management, cybersecurity, trading, and gaming. The model performs supervised next-action prediction conditioned on historical observations, actions, rewards, and termination signals. For the first time, it demonstrates that a single Transformer policy can match the performance of task-specific policies across over a thousand heterogeneous tasks, establishing a unified paradigm for large-scale multitask reinforcement learning.