π€ AI Summary
DreamerV3 exhibits limited long-term memory and long-horizon dependency modeling capabilities in complex environments such as Atari-Boxing, Freeway, Pong, and Crafter. To address this, we propose TransDreamerV3βthe first integration of a full-attention Transformer encoder throughout the entire DreamerV3 world model pipeline, replacing its recurrent sequence modeling with global temporal modeling. Specifically, the Transformer is incorporated into all core components: RSSM-based dynamics modeling, Actor-Critic policy optimization, and self-supervised reconstruction. This unified architectural redesign enhances robustness in sparse-reward, long-horizon decision-making. Empirical results demonstrate that TransDreamerV3 outperforms the original DreamerV3 on Atari-Freeway and Crafter, validating the critical role of Transformer-enhanced temporal representation learning in world models. The approach establishes a principled framework for end-to-end, attention-driven sequential reasoning in model-based reinforcement learning.
π Abstract
This paper introduces TransDreamerV3, a reinforcement learning model that enhances the DreamerV3 architecture by integrating a transformer encoder. The model is designed to improve memory and decision-making capabilities in complex environments. We conducted experiments on Atari-Boxing, Atari-Freeway, Atari-Pong, and Crafter tasks, where TransDreamerV3 demonstrated improved performance over DreamerV3, particularly in the Atari-Freeway and Crafter tasks. While issues in the Minecraft task and limited training across all tasks were noted, TransDreamerV3 displays advancement in world model-based reinforcement learning, leveraging transformer architectures.