decision transformer modeling

Designs, implements, and evaluates Transformer-based sequence models that ingest past states, actions, and conditioning signals (e.g., returns) to predict or generate future actions and state-action trajectories; builds end-to-end pipelines to extract policies from offline trajectories using the Decision Transformer paradigm. Analyzes model behavior for long-term dependency capture, consistency between behavioral and environmental dynamics, and the fidelity of generated sequences for downstream decision-making.

decisiontransformermodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the input redundancy inherent in the Decision Transformer (DT) when utilizing Return-to-Go (RTG) sequences, which compromises both computational efficiency and performance. The authors propose the Decoupled Decision Transformer (DDT), which is the first to explicitly identify the redundancy in RTG sequences and decouple the RTG conditioning mechanism: only the most recent RTG value is used to guide action prediction, while the Transformer backbone processes solely the observation and action sequences. This streamlined architecture reduces unnecessary computation, enhances inference efficiency, and achieves significant performance improvements over the original DT across multiple offline reinforcement learning benchmarks, matching or surpassing the performance of current state-of-the-art DT variants.

Decision Transformeroffline reinforcement learningredundancy

This work addresses the inefficiency of the standard Decision Transformer, which embeds Return-to-Go (RTG) as standalone tokens in the autoregressive sequence, resulting in unnecessarily long sequences and high computational overhead. The authors propose decoupling the RTG conditioning from the autoregressive sequence and instead injecting it directly into the state representations. This enables the model to operate solely on compact (state, action) sequences, achieving a novel separation between sparse RTG signals and dense trajectory information. Evaluated on the D4RL benchmark, the proposed method significantly outperforms the standard Decision Transformer, attaining state-of-the-art performance while reducing sequence length by approximately one-third and substantially improving inference efficiency.

computational efficiencyDecision Transformeroffline reinforcement learning

Predictive Coding for Decision Transformer

Oct 04, 2024
TL
T. Luu
🏛️ KAIST

Decision Transformers (DTs) underperform in sparse-reward, long-horizon offline goal-directed reinforcement learning, primarily because scalar return conditioning fails to capture temporal compositionality. To address this, we propose Predictive Coding Conditioning (PCC), a novel conditioning mechanism that replaces scalar returns with generalized future-state or goal representations—marking the first integration of predictive coding into the DT framework. PCC enables joint policy modeling conditioned on both past observations and future goals. Our method extends the Transformer architecture with predictive coding–based representations, future-conditioned modeling, and offline behavior cloning training. Evaluated on eight challenging offline benchmarks across AntMaze and FrankaKitchen, PCC matches or surpasses state-of-the-art value-based and Transformer-based baselines. Furthermore, it demonstrates strong generalization in real-world physical robot goal-reaching tasks.

Addresses inefficiency in return conditioning methodsEnhances performance in sparse-reward long-horizon tasksImproves decision transformers in offline RL tasks

Existing offline safe reinforcement learning methods struggle with complex, multi-threaded, and temporally dependent real-world constraints. This paper proposes STL-Decision Transformer, the first approach to explicitly incorporate Signal Temporal Logic (STL) specifications as conditional inputs into the Decision Transformer architecture, enabling joint optimization of reward maximization and satisfaction of multi-granularity temporal safety constraints. By directly encoding STL semantics into the policy conditioning mechanism, our method overcomes fundamental limitations of conventional conditional policies in expressivity and generalizability for temporal logic, while supporting continuous, controllable adjustment of STL satisfaction degrees. Evaluated on the DSRL benchmark, STL-Decision Transformer significantly outperforms state-of-the-art methods, achieving simultaneous improvements in both cumulative reward and constraint satisfaction rate. These results demonstrate its effectiveness and robustness for offline, constraint-driven policy learning under rich temporal safety requirements.

Complex Multi-threaded ProblemsOffline Reinforcement LearningSequential Dependence

State-Space Modeling in Long Sequence Processing: A Survey on Recurrence in the Transformer Era

Jun 13, 2024
MT
Matteo Tiezzi
🏛️ IIT | University of Siena | IMT

Long-sequence modeling faces fundamental challenges including limited context length, difficulty in capturing long-range dependencies, and low efficiency in online learning. To address these, this work systematically reviews the resurgence of state-space models (SSMs) and recurrent computation, proposing a novel local forward-computation paradigm tailored for real-world online learning—thereby circumventing the temporal backtracking constraints inherent in standard backpropagation through time (BPTT). We introduce the first unified taxonomy encompassing both deep SSMs and large-context Transformers. Our framework integrates structured linear attention, enhanced RNN architectures, local recurrence mechanisms, and online optimization algorithms. The study rigorously clarifies the theoretical representational advantages and practical sequential reasoning benefits of recurrent modeling over alternatives. Collectively, this work delivers a scalable technical roadmap for low-latency, highly extensible long-sequence modeling.

Addressing limitations of Transformers with state-space modelsExploring efficient online learning beyond backpropagation through timeSurveying recurrent models for long sequence processing

Latest Papers

What's happening recently
View more

A Comparison Between Decision Transformers and Traditional Offline Reinforcement Learning Algorithms

Nov 20, 2025
AM
Ali Murtaza Caunhye
🏛️ University of KwaZulu-Natal | Centre for Artificial Intelligence Research

This study investigates how decision transformers (DTs) compare to conventional offline reinforcement learning (RL) algorithms—specifically conservative Q-learning (CQL) and implicit Q-learning (IQL)—under varying reward densities (dense vs. sparse) in the ANT continuous-control benchmark. Method: We conduct a systematic, controlled empirical evaluation across uniformly configured offline datasets of varying quality and reward sparsity. Contribution/Results: We find that DTs exhibit remarkable robustness to reward density shifts: they outperform both CQL and IQL in sparse-reward regimes and under medium-quality offline data, achieving higher policy performance, greater stability, and lower evaluation variance. In contrast, IQL excels in dense-reward settings, while CQL demonstrates superior overall robustness across diverse conditions. Crucially, this work provides the first empirical evidence that sequence modeling—via autoregressive action prediction—confers distinct advantages in low signal-to-noise-ratio feedback environments. These findings offer principled guidance for reward-structure-aware algorithm selection and design in offline RL.

Analyzes algorithm sensitivity to reward density and data qualityCompares Decision Transformers with traditional offline RL algorithmsEvaluates performance in dense versus sparse reward environments

This work addresses the limitations of traditional reinforcement learning in handling long-horizon sequential decision-making within complex dynamic environments—such as real-time bidding in computational advertising—and the difficulty of large language models (LLMs) in modeling continuous numerical values. To overcome these challenges, the authors propose DecisionLLM, a novel framework that treats trajectory data as an independent modality and aligns it with natural language task descriptions, enabling LLMs to autoregressively predict future decisions. This approach marks the first effective application of LLMs to offline long-sequence decision tasks. Experimental results demonstrate that DecisionLLM-3B outperforms Decision Transformer by 69.4 on Maze2D umaze-v1 and by 0.085 on the AuctionNet benchmark, while also revealing clear scaling laws with respect to model size, data volume, and data quality.

Continuous valuesLarge Language ModelsLong-sequence decision-making

Unified token representations for sequential decision models

Oct 24, 2025
ZT
Zhuojing Tian
🏛️ Intelligent Game and Decision Lab(IGDL) | Tsinghua University

Existing offline reinforcement learning methods—such as Decision Transformers—are hindered by redundant tokenization and quadratic attention complexity, compromising both computational efficiency and generalization. To address this, we propose Unified Token Representation (UTR), which encodes return, state, and action into a single compact token, drastically shortening sequence length and reducing computational overhead. Theoretically, UTR yields a tighter Rademacher complexity bound, thereby enhancing generalization. Leveraging UTR, we design two architectures: UDT (a Transformer-based variant) and UDC (a gated CNN-based variant). On standard offline RL benchmarks, both achieve state-of-the-art or competitive performance while accelerating inference by 2–5× and reducing memory footprint by 30%–60%. Moreover, UTR demonstrates strong cross-architecture transfer robustness.

Addresses quadratic attention complexity in offline RLImproves scalability for resource-constrained control systemsReduces redundant tokenization in decision transformers

This work addresses the challenge of achieving unified and scalable multitask decision-making across a vast array of heterogeneous reinforcement learning environments. It introduces LDM-v0, a Transformer-based universal policy model trained offline at scale on multimodal trajectory data from approximately 1,000 diverse domains—including robotics, autonomous driving, inventory management, cybersecurity, trading, and gaming. The model performs supervised next-action prediction conditioned on historical observations, actions, rewards, and termination signals. For the first time, it demonstrates that a single Transformer policy can match the performance of task-specific policies across over a thousand heterogeneous tasks, establishing a unified paradigm for large-scale multitask reinforcement learning.

heterogeneous environmentslarge decision modelsmulti-task reinforcement learning

This work addresses the limitations of standard Decision Transformers in robotic manipulation, which suffer from low sample efficiency, insufficient exploration, and suboptimal performance due to reliance on uniform experience replay. To overcome these issues, the authors propose the E²DT framework, which innovatively integrates k-Determinantal Point Processes (k-DPPs) with the Decision Transformer to enable experience-aware sampling through a joint quality-diversity kernel. Trajectory diversity is quantified via latent embeddings, while trajectory quality is assessed by combining return-to-go (RTG) quantiles, predictive uncertainty, and stage coverage. Experimental results demonstrate that E²DT significantly outperforms existing methods in both simulated and real-world robotic tasks, markedly improving sample efficiency and robustness in long-horizon reinforcement learning settings.

Decision Transformerexperience samplingexploration-exploitation trade-off

Hot Scholars

JB

Jiang Bian

Regenstrief Institue; Indiana University; IU Health
data sciencereal-world dataontology/semanticeHealth/social media
ME

Melike Erol-Kantarci

Canada Research Chair & Professor, University of Ottawa and Sr. Product Manager for AI RAN, Ericsson
AI-enabled wireless networksAIGenAI5G6GO-RANsmart grid
XT

Xin Tao

Kuaishou
Computer VisionGenerative AI
ZL

Zihao Li

China University of Geoscience, Wuhan
Computer VisionRemote SensingDeep Learning
HL

Haoxuan Li

Peking University
causal inferencerecommender systemtrustworthy AIlarge language model