TransDreamerV3: Implanting Transformer In DreamerV3

πŸ“… 2025-06-20
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
DreamerV3 exhibits limited long-term memory and long-horizon dependency modeling capabilities in complex environments such as Atari-Boxing, Freeway, Pong, and Crafter. To address this, we propose TransDreamerV3β€”the first integration of a full-attention Transformer encoder throughout the entire DreamerV3 world model pipeline, replacing its recurrent sequence modeling with global temporal modeling. Specifically, the Transformer is incorporated into all core components: RSSM-based dynamics modeling, Actor-Critic policy optimization, and self-supervised reconstruction. This unified architectural redesign enhances robustness in sparse-reward, long-horizon decision-making. Empirical results demonstrate that TransDreamerV3 outperforms the original DreamerV3 on Atari-Freeway and Crafter, validating the critical role of Transformer-enhanced temporal representation learning in world models. The approach establishes a principled framework for end-to-end, attention-driven sequential reasoning in model-based reinforcement learning.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Agent Architectures

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
πŸ“ Abstract
This paper introduces TransDreamerV3, a reinforcement learning model that enhances the DreamerV3 architecture by integrating a transformer encoder. The model is designed to improve memory and decision-making capabilities in complex environments. We conducted experiments on Atari-Boxing, Atari-Freeway, Atari-Pong, and Crafter tasks, where TransDreamerV3 demonstrated improved performance over DreamerV3, particularly in the Atari-Freeway and Crafter tasks. While issues in the Minecraft task and limited training across all tasks were noted, TransDreamerV3 displays advancement in world model-based reinforcement learning, leveraging transformer architectures.
Problem

Research questions and friction points this paper is trying to address.

Enhancing DreamerV3 with transformer for better RL performance
Improving memory and decision-making in complex environments
Advancing world model-based RL via transformer architecture
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates transformer encoder into DreamerV3
Enhances memory and decision-making in RL
Improves performance in complex environments
πŸ”Ž Similar Papers
No similar papers found.