DHP: Discrete Hierarchical Planning for Hierarchical Reinforcement Learning Agents

📅 2025-02-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the limitations of conventional distance-based methods in long-horizon visual planning—specifically their inability to model long-range dependencies and enable efficient re-planning within hierarchical reinforcement learning (HRL)—this paper proposes Discrete Hierarchical Planning (DHP). DHP achieves end-to-end hierarchical decision optimization via recursive generation of discrete abstract subgoals, tree-structured trajectory advantage estimation (which implicitly favors short-horizon plans while enabling ultra-deep generalization), and on-policy imagined-data-driven SAC training with active exploration. Key innovations include: (i) the first discrete subgoal planning paradigm for visual HRL; (ii) a tree-aware advantage estimator; and (iii) a closed-loop imagination–exploration co-training mechanism. Evaluated on a 25-room long-horizon visual navigation task, DHP significantly improves success rate, reduces average episode length, and lowers planning complexity to O(log N). Ablation studies confirm the critical contribution of each component.

Technology Category

Humans and AI: Human-Aware Planning and Behavior PredictionPlanning, Routing, and Scheduling: Mixed Discrete/Continuous PlanningMultiagent Systems: Multiagent Planning

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingResponsible Web: Human-perceived consequences of algorithmic deployment on the webGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphs
📝 Abstract
In this paper, we address the challenge of long-horizon visual planning tasks using Hierarchical Reinforcement Learning (HRL). Our key contribution is a Discrete Hierarchical Planning (DHP) method, an alternative to traditional distance-based approaches. We provide theoretical foundations for the method and demonstrate its effectiveness through extensive empirical evaluations. Our agent recursively predicts subgoals in the context of a long-term goal and receives discrete rewards for constructing plans as compositions of abstract actions. The method introduces a novel advantage estimation strategy for tree trajectories, which inherently encourages shorter plans and enables generalization beyond the maximum tree depth. The learned policy function allows the agent to plan efficiently, requiring only $log N$ computational steps, making re-planning highly efficient. The agent, based on a soft-actor critic (SAC) framework, is trained using on-policy imagination data. Additionally, we propose a novel exploration strategy that enables the agent to generate relevant training examples for the planning modules. We evaluate our method on long-horizon visual planning tasks in a 25-room environment, where it significantly outperforms previous benchmarks at success rate and average episode length. Furthermore, an ablation study highlights the individual contributions of key modules to the overall performance.
Problem

Research questions and friction points this paper is trying to address.

Long-horizon visual planning tasks
Hierarchical Reinforcement Learning
Discrete Hierarchical Planning method
Innovation

Methods, ideas, or system contributions that make the work stand out.

Discrete Hierarchical Planning method
Novel advantage estimation strategy
On-policy imagination data training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shashank Sharma
Department of Computer Science, University of Bath
J
Janina Hoffmann
Department of Psychology, University of Bath
V
Vinay Namboodiri
Department of Computer Science, University of Bath