🤖 AI Summary
This study addresses the option explosion problem in existing option discovery algorithms, which arises from their reliance on full-state goals and impedes exploration and reward maximization in sparse-reward environments. We propose a generalized option construction method based on learning subsets of abstract features, overcoming the limitations of full-state matching. By integrating intrinsic motivation-based reinforcement learning, temporal abstraction, and image feature selection techniques, our approach generates transferable abstract options. This method significantly reduces the number of required options while enhancing generalization capability. Empirical evaluations demonstrate that it enables efficient exploration across three image-based sparse-reward tasks, including Montezuma's Revenge.
📝 Abstract
Temporal abstraction via options can improve exploration in large environments. However, existing option discovery algorithms find subgoals that target all aspects of the state simultaneously. This state-reaching approach produces options that only apply in narrow regions of the state-space, eventually causing an explosion in the number of options that overwhelms the agent, and impedes progress on its primary task of reward maximization. We introduce an algorithm that instead identifies a small, relevant subset of features for each subgoal, yielding options that generalize broadly and accelerate exploration. Our approach learns abstract, transferrable options and achieves rapid exploration in three sparse-reward, image-based domains, including the Atari game MontezumasRevenge.