🤖 AI Summary
This work addresses the challenges of state representation and poor sample efficiency in visual reinforcement learning caused by high-dimensional image inputs. The authors propose a self-supervised auxiliary task based on masked prediction, which leverages observation sequences and their contextual information collected by the agent to learn compact yet informative sequential representations in a latent space using a Transformer architecture. Unlike conventional reconstruction-based approaches, this method employs a non-reconstructive masked prediction mechanism that emphasizes understanding of task dynamics rather than pixel-level fidelity. Experimental results demonstrate that the proposed approach significantly outperforms state-of-the-art sample-efficient reinforcement learning algorithms across multiple continuous and discrete control benchmarks, achieving substantial gains in sample efficiency.
📝 Abstract
Vision-based deep reinforcement learning involves dealing with high-dimensional inputs of image information. It is crucial to abstract effective states from high-dimensional image inputs and limited samples for sample-efficient reinforcement learning. To address this challenge, inspired by fields such as natural language processing and computer vision, we propose a self-supervised task based on mask prediction as an auxiliary task for reinforcement learning. This non-reconstruction method uses the sequence information collected by the agent from the environment and the context information in the sequence to predict the masked information, thereby strengthening the agent's understanding of the task and learning effective representations. Combined with transformers, we find that the model reconstructs the masked input sequence in the latent space. By feeding the compressed representations learned by this method into reinforcement learning models, we observe an improvement in the sample efficiency of reinforcement learning. Moreover, the model outperforms state-of-the-art sample-efficient reinforcement learning methods on multiple continuous and discrete control benchmarks.