hierarchical belief-based planning

Design and implement planners that reason over belief states using hierarchical abstractions and macro-actions to produce multi-step sequential plans. Build rollout and evaluation mechanisms (e.g., finite-horizon rollouts on hierarchical structures such as an HSG) that fuse prior knowledge with online evidence to estimate long-term expected returns and select globally consistent actions that reduce backtracking.

hierarchicalbelief-basedplanning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.37
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$197K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenges of large action spaces, high stochasticity, long decision horizons, and resource constraints in sequential stochastic combinatorial optimization by proposing a model-based hierarchical reinforcement learning framework. The approach integrates a world model formulated as a semi-Markov decision process (SMDP) with a latent-space tree-search planner, constructing temporal structures of abstract actions through multi-timescale objectives and jointly learning budget-aware policies conditioned on subgoals to enable context-sensitive resource allocation. A key innovation lies in incorporating adaptive temporal abstraction into hierarchical planning, allowing the latent dynamics to explicitly capture the effective duration of abstract actions. Experimental results demonstrate that the proposed method significantly outperforms strong existing baselines across multiple challenging benchmarks.

Hierarchical Reinforcement LearningMulti-Timescale AbstractionSemi-Markov Decision Process

In multi-objective sequential decision-making, conventional goal-conditioned (GC) policies—optimized solely for the current goal—often cause path blockages and render subsequent goals unreachable due to neglect of inter-goal dependencies. Method: We propose Dual-Objective Conditional MDPs and Two-Step Lookahead Goal-Conditioned MDPs, the first frameworks to explicitly model sequential goal dependencies within the GC policy’s conditioning mechanism. Built upon the TD3+HER architecture, our approach establishes a novel hierarchical RL paradigm that jointly optimizes for both the current and subsequent goal reachability during policy training. Contribution/Results: Evaluated on navigation and inverted pendulum tasks, our method significantly improves policy stability and sample efficiency, consistently outperforming standard GC-MDPs and single-objective GC baselines. It effectively overcomes the local optimization limitation inherent in traditional GC policies by incorporating lookahead goal dependencies into the conditioning structure.

Addresses failure in multi-goal planning when intermediate goals block subsequent goalsImproves stability and efficiency by conditioning on next two goalsProposes MDPs optimizing for current and future goals simultaneously

DHP: Discrete Hierarchical Planning for Hierarchical Reinforcement Learning Agents

Feb 04, 2025
SS
Shashank Sharma
🏛️ University of Bath

To address the limitations of conventional distance-based methods in long-horizon visual planning—specifically their inability to model long-range dependencies and enable efficient re-planning within hierarchical reinforcement learning (HRL)—this paper proposes Discrete Hierarchical Planning (DHP). DHP achieves end-to-end hierarchical decision optimization via recursive generation of discrete abstract subgoals, tree-structured trajectory advantage estimation (which implicitly favors short-horizon plans while enabling ultra-deep generalization), and on-policy imagined-data-driven SAC training with active exploration. Key innovations include: (i) the first discrete subgoal planning paradigm for visual HRL; (ii) a tree-aware advantage estimator; and (iii) a closed-loop imagination–exploration co-training mechanism. Evaluated on a 25-room long-horizon visual navigation task, DHP significantly improves success rate, reduces average episode length, and lowers planning complexity to O(log N). Ablation studies confirm the critical contribution of each component.

Discrete Hierarchical Planning methodHierarchical Reinforcement LearningLong-horizon visual planning tasks

This work addresses the limitations of existing large language model (LLM) agents, which typically employ fixed-granularity planning mechanisms that struggle to balance efficiency on simple tasks with the detailed reasoning required for complex ones. To overcome this, the paper introduces AdaPlan-H, a cognitively inspired adaptive hierarchical planning framework that, for the first time, integrates a progressive refinement strategy into LLM agents to enable dynamic, task-difficulty-aware adjustment of planning granularity. By synergistically combining hierarchical task decomposition, imitation learning, and capability enhancement, AdaPlan-H supports the adaptive generation and continuous optimization of planning hierarchies. Experimental results demonstrate that AdaPlan-H significantly improves success rates on multi-step complex tasks while effectively avoiding over-planning, thereby validating its efficiency and flexibility.

adaptive granularityhierarchical planningLLM agents

HyperTree Planning: Enhancing LLM Reasoning via Hierarchical Thinking

May 05, 2025
RG
Runquan Gui
🏛️ University of Science and Technology of China | Huawei Technologies | Tianjin University

To address challenges in complex planning tasks—including excessively long reasoning chains, diverse constraints, and heterogeneous subtasks—this paper proposes the hypertree-structured planning paradigm. Methodologically, it adopts a divide-and-conquer strategy to enable hierarchical and dynamic reasoning: (i) it introduces the hypertree structure to model multi-granularity planning processes with constraint awareness and scalable inference; (ii) it designs an autonomous iterative refinement framework that overcomes expressivity limitations of conventional linear or tree-structured planners; and (iii) it develops an LLM-based dynamic hypertree expansion mechanism, a constraint-driven subtask decomposition and coordinated scheduling algorithm, and an iterative planning outline optimization technique. Evaluated on the TravelPlanner benchmark, our approach achieves state-of-the-art accuracy using Gemini-1.5-Pro, outperforming o1-preview by 3.6× in planning performance.

Addresses complex planning challenges in LLMsEnhances reasoning via hierarchical hypertree structuresImproves performance in multi-subtask constraint handling

Latest Papers

What's happening recently
View more

Offline goal-conditioned reinforcement learning often suffers from low sample efficiency due to redundancy in state-goal pairs. This work proposes a hierarchical policy that abstracts goal conditioning away from an absolute reference frame by introducing relativized options and multi-level state representations, enabling experience reuse across similar contexts. Unlike conventional hierarchical approaches that primarily provide temporal abstraction, the proposed architecture explicitly exploits hierarchy to eliminate redundancy in the state-goal space. Experimental results demonstrate that the method substantially improves performance on offline goal-conditioned tasks, validating the effectiveness of the introduced abstract inductive bias.

AbstractionGoal-Conditioned Reinforcement LearningMarkov Decision Processes

Simulation-based planning with rollouts is a widely-deployed technique for decision making in stochastic environments. The primary instrument of simulation-based planning is a sampling model, which is repeatedly called to generate trajectories and estimate the utilities of available actions. Among the actions thus explored, one with the maximum estimated utility is then executed. In this paper, we examine the effect of using common random numbers in the simulation process. We obtain a simple recipe for (provably) reducing variance in relative utility when simulations invoke a rollout policy beyond some depth. Experiments on synthetic tasks confirm that our scheme improves task performance. The broader significance of our innovation is apparent from two practical applications: (1) single-step lookahead planning in a pension-disbursement task, and (2) a deployment of the well-known UCT algorithm for the game of Ludo.

common random numbersrolloutssimulation-based planning

This work addresses the critical yet poorly understood role of long-horizon multi-turn planning in foundation model agents, which is hindered by the uncontrolled nature of internet-scale pretraining data. The authors construct a unified and controllable multi-turn environment to systematically investigate how the format, distribution, and quality of pretraining data influence planning capabilities. During post-training, they introduce GRPO, Online Policy Distillation (OPD), and Multi-teacher Online Policy Distillation (MOPD) to shape and integrate planning skills. Their findings reveal the essential role of explicit world models in enabling long-horizon generalization, demonstrate OPD’s superiority over GRPO under low-quality long-horizon data, and present the first successful fusion and transfer of cross-environment planning abilities via MOPD. Experiments further validate the efficacy of limited high-quality data and elucidate MOPD’s robust generalization, continual learning, and interference resilience across compatible, partially shared, or conflicting planning paradigms.

foundation model agentslong-horizon planningmulti-turn planning

This work addresses the limited policy expressivity and insufficient exploration efficiency in hierarchical reinforcement learning by proposing a novel paradigm in which a controller constructs a flexible behavioral space through linear combinations of multiple option reward functions. Departing from conventional assumptions of long-horizon planning, the approach demonstrates that the core advantage of hierarchical structures lies in their capacity to enhance exploration. Empirical evaluation in the NetHack Learning Environment shows that the proposed method substantially improves both exploration efficiency and overall learning performance, thereby validating its effectiveness and superiority in complex environments.

behaviour spacesexplorationhierarchical reinforcement learning

Hot Scholars

JZ

Jingren Zhou

Alibaba Group, Microsoft
Cloud ComputingLarge Scale Distributed SystemsMachine LearningQuery Processing
JH

Jianye Hao

Huawei Noah's Ark Lab/Tianjin University
Multiagent SystemsEmbodied AI
GJ

Gregory J. Stein

Assistant Professor, George Mason University
machine learningroboticsplanning under uncertaintynavigation
AK

Alois Knoll

Technische Universität München
RoboticsAISensor Data FusionAutonomous Driving