Direction-Conditioned Policies for Online Goal-Conditioned Reinforcement Learning

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation in contrastive reinforcement learning where policies fail to effectively exploit the geometric structure of the representation space. To overcome this, we propose a direction-conditioned policy that, for the first time, incorporates both the direction and distance within the representation space as policy conditions. During training, visited states serve as waypoints to guide exploration, while at deployment, the agent directly points toward the goal. This approach leverages geometric priors without explicit planning, unifying the interfaces for training and deployment. Empirical evaluations across nine tasks demonstrate significant improvements in success rate and goal proximity, with the policy more accurately capturing shortest-path geometric characteristics. Furthermore, we validate the causal influence of directional information on policy behavior, endowing the model with zero-shot planning capabilities.
📝 Abstract
Contrastive Reinforcement Learning (CRL) learns representations that estimate goal reachability, yet its policy remains conditioned on raw goals and therefore does not directly exploit the geometry encoded by its critic. We introduce Direction-Conditioned Policies (DCP), a method built around a small modification to CRL: DCP selects previously visited states as waypoints during online training and conditions the policy on their direction and distance in representation space. At deployment, DCP applies the same interface directly to the final goal, requiring neither waypoint selection nor planning. Across nine navigation and manipulation tasks, DCP attains higher final success rates than CRL on seven tasks and spends more time near the goal on seven. Controlled maze experiments further show that DCP captures shortest-path geometry more accurately and that the supplied direction causally influences the actor's behavior. We identify waypoint coverage and ranking as limits to exploration, and show that learned candidate generation improves goal reaching in two controlled mazes.
Problem

Research questions and friction points this paper is trying to address.

Contrastive Reinforcement Learning
Goal-Conditioned Reinforcement Learning
Representation Geometry
Exploration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Direction-Conditioned Policies
Contrastive Reinforcement Learning
Goal-Conditioned Reinforcement Learning
Representation Space
Waypoint Selection
🔎 Similar Papers
No similar papers found.
S
S K Swaminathan
Department of Computer Science and Engineering, Indian Institute of Technology Kharagpur, India
D
Damiya Gondha
Department of Computer Science and Engineering, Indian Institute of Technology Kharagpur, India
T
Theyanesh Eswaramoorthy Rajahkrishnan
Department of Mechanical Engineering, Indian Institute of Technology Kharagpur, India
Aritra Hazra
Aritra Hazra
Department of Computer Science and Engineering, IIT Kharagpur
Formal MethodsDesign VerificationCAD for SecurityArtificial IntelligenceMachine Learning