🤖 AI Summary
This study addresses the limitation in contrastive reinforcement learning where policies fail to effectively exploit the geometric structure of the representation space. To overcome this, we propose a direction-conditioned policy that, for the first time, incorporates both the direction and distance within the representation space as policy conditions. During training, visited states serve as waypoints to guide exploration, while at deployment, the agent directly points toward the goal. This approach leverages geometric priors without explicit planning, unifying the interfaces for training and deployment. Empirical evaluations across nine tasks demonstrate significant improvements in success rate and goal proximity, with the policy more accurately capturing shortest-path geometric characteristics. Furthermore, we validate the causal influence of directional information on policy behavior, endowing the model with zero-shot planning capabilities.
📝 Abstract
Contrastive Reinforcement Learning (CRL) learns representations that estimate goal reachability, yet its policy remains conditioned on raw goals and therefore does not directly exploit the geometry encoded by its critic. We introduce Direction-Conditioned Policies (DCP), a method built around a small modification to CRL: DCP selects previously visited states as waypoints during online training and conditions the policy on their direction and distance in representation space. At deployment, DCP applies the same interface directly to the final goal, requiring neither waypoint selection nor planning. Across nine navigation and manipulation tasks, DCP attains higher final success rates than CRL on seven tasks and spends more time near the goal on seven. Controlled maze experiments further show that DCP captures shortest-path geometry more accurately and that the supplied direction causally influences the actor's behavior. We identify waypoint coverage and ranking as limits to exploration, and show that learned candidate generation improves goal reaching in two controlled mazes.