🤖 AI Summary
This study addresses the reliance of intrinsic motivation in unsupervised reinforcement learning on manually selected informative variables and domain expertise. To overcome this limitation, it proposes the Forward Controllable Information Production (F-CIP) objective, formulating it for the first time as a native reinforcement learning expression defined solely by system dynamics without requiring variable selection. By integrating the F-CIP algorithm, an unsupervised reinforcement learning framework is constructed that entirely eliminates dependence on manual reward engineering and informative variable selection. Through this approach, agents autonomously discover fundamental behaviors such as balancing and can acquire coordinated gaits, including jumping and running, using only a simple velocity reward.
📝 Abstract
Reinforcement learning (RL) is a powerful paradigm for training agents, yet its success rests on domain expertise of human engineers who design informative reward signals for every new task. Unsupervised RL aims to reduce this engineering with intrinsic motivation (IM): reward signals that emerge from the agent environment interaction itself. Existing IM objectives, however, involve the selection of information variables, which re-introduces domain expertise the field has sought to eliminate. We introduce Forward CIP (F-CIP), an RL-native formulation of the Controllable Information Production (CIP) objective, which is defined by the system's dynamics alone and requires no such selection. We prove that F-CIP is compatible with RL and demonstrate its effectiveness with existing algorithms. Training agents with F-CIP results in unsupervised discovery of primitive behaviors such as balancing and maintaining controllability, which are essential for more complex robot behaviors. Paired with a simple forward-velocity reward, our method produces coordinated gaits such as hopping and running which otherwise require reward engineering to learn.