Score
Designs, implements, and evaluates heuristic functions and the search strategies that use them to guide state‑space or motion planners (e.g., A* and related algorithms). This includes creating admissible and informative heuristics, designing search‑space representations and pruning/ranking/filtering rules, integrating cost or surrogate terms, tuning heuristic parameters, and analyzing the tradeoffs between solution optimality and search efficiency.
Handcrafted heuristic functions in search-based navigation suffer from poor generalization across unseen maps and long-distance paths. Method: This paper proposes a local heuristic learning framework that explicitly defines and end-to-end learns either heuristic bias correction or local cost estimation within a spatial neighborhood—replacing conventional global heuristic modeling. By decomposing complex global prediction into lightweight local regression, the approach significantly reduces learning complexity. Integrated with graph search algorithms (e.g., A*), it operates under supervised learning using local state inputs while preserving bounded suboptimality guarantees. Contribution/Results: Experiments demonstrate 2–20× reduction in node expansions, improved training efficiency, and robust generalization to both unseen maps and long-range trajectories—without compromising solution quality or theoretical guarantees.
Dynamic heuristics—functions updated in real time based on search history—pose a fundamental challenge to optimality guarantees in heuristic search. Traditional analyses assume static heuristics, rendering existing theoretical foundations inapplicable. Method: This work formally defines dynamic heuristics and introduces a generalized search framework that integrates dynamism into the foundational design of A*-style algorithms. Through formal semantic modeling and rigorous analysis of dynamic heuristic functions, we establish sufficient conditions for admissibility and optimality. Contribution/Results: We prove that our algorithm variants retain admissibility and guarantee optimal solutions under any valid dynamic heuristic. Our framework transcends the static-heuristic assumption, unifying and explaining classical planning heuristics—including FF and LM-cut—under a common theoretical umbrella. This provides the first rigorous, provable foundation for dynamic heuristic search, advancing it from empirical practice to a sound, analyzable algorithmic paradigm.
This work addresses the weak generalization of neural heuristics under sample scarcity in classical planning. It investigates how sample generation strategies affect the performance of Greedy Best-First Search (GBFS). We propose three sampling strategies to enhance state representativeness and two modeling techniques to improve cost estimate accuracy, uncovering a tight coupling between sample distribution quality and heuristic estimation error. Controlled experiments demonstrate that, under limited training data, GBFS guided by our learned neural heuristic achieves over 30% higher average solution coverage than baseline methods, significantly improving both few-shot generalization and search efficiency. Our core contribution is establishing an interpretable, causal linkage among sample generation, estimation accuracy, and search performance—yielding a practical, reproducible framework for neural heuristic learning in resource-constrained planning settings.
This work addresses time-dependent planning in fixed domains, where conventional planners suffer from heuristic cold-start and MDP state truncation under limited training problem sets. To overcome these bottlenecks, we propose a novel symbolic-heuristic-guided reinforcement learning paradigm: (i) designing RL reward functions grounded in symbolic heuristics (e.g., h_max, LM-cut); (ii) introducing a residual heuristic learning framework to mitigate initial policy bias; and (iii) constructing a hybrid search architecture with multi-priority queues to enhance search guidance. Our method integrates DQN/PPO variants with symbolic domain knowledge, enabling joint planning-and-learning modeling. Evaluated across multiple IPC temporal domains, it achieves a 37% improvement in both solution success rate and runtime over baselines, establishing new state-of-the-art performance. The approach offers an interpretable and transferable pathway for heuristic synthesis in automated planning.
This work addresses the limitations of existing large language model (LLM)-driven approaches to automated heuristic design, which are constrained by fixed evolutionary rules and static prompting templates, hindering long-horizon reasoning and efficient evolution. The authors propose modeling heuristic generation as a sequential decision-making process over an entailment graph, introducing the entailment graph as a stateful memory structure that enables cross-generation information reuse and conflict avoidance. A multi-agent collaborative framework is developed, comprising a policy agent that plans evolutionary actions, a world model agent that simulates heuristic performance, and a critic agent that performs routing-based reflection, thereby transforming trial-and-error evolution into state-aware, planning-driven search. The method achieves significantly faster convergence and yields superior heuristics across multiple combinatorial optimization problems, demonstrates compatibility with diverse LLM backbones, and exhibits strong scalability.
Traditional heuristics are computationally intractable in numeric planning domains with infinite action spaces, primarily due to unbounded control parameters affecting action effects and preconditions. Method: We propose an optimistic compilation framework that abstracts parameter-dependent action effects into bounded constant effects and relaxes preconditions, thereby transforming the infinite action space into a finite, tractable one. This enables the first successful extension of subgoal-based heuristics to infinite-action settings via structured abstraction, reusing classical heuristic patterns. Technically, our approach integrates numeric planning relaxation, control-parameter abstraction, and subgoal decomposition. Contribution/Results: Evaluated on multiple benchmark domains, our method renders classical heuristics computationally feasible for the first time in such infinite-action numeric planning problems. It significantly outperforms existing baselines in both solving efficiency and scalability, demonstrating robustness and practical applicability.
This work proposes A-CEoH, a domain-agnostic method for automatically generating high-quality heuristic functions for A* search without relying on manual expert design. By embedding the A* algorithm code directly into the prompting context and leveraging the in-context learning capabilities of large language models (LLMs) within an Evolutionary framework for Heuristics (EoH), A-CEoH autonomously constructs effective heuristics with no human intervention. Evaluated on benchmark domains such as the Uncertain Partially Observable Multi-agent Pathfinding Problem (UPMP) and sliding tile puzzles, the heuristics produced by A-CEoH significantly outperform existing baselines and even surpass handcrafted expert-designed heuristics, thereby overcoming the longstanding dependency of traditional A* search on human-crafted domain knowledge.
This work proposes an intention-driven heuristic approach to enhance the efficiency of classical planning. Inspired by intention modeling in goal recognition, the authors introduce—for the first time—a reversal of trajectory-to-goal directedness evaluation into the planning domain, constructing a novel heuristic function to guide search. By integrating probabilistic intention inference with classical planning, they design a computationally efficient heuristic evaluation framework. Two new heuristics derived from this framework have been incorporated into state-of-the-art planners and demonstrate significant performance improvements across multiple benchmark domains, thereby validating the effectiveness of leveraging goal recognition perspectives to empower classical planning.
This work addresses the long-standing challenge that reinforcement learning-based hyper-heuristics (RLHH) often struggle to effectively select low-level heuristics on standard benchmark functions. Focusing on the LeadingOnes problem and combining two randomized local search operators—RLS₁ and RLS₂—the study provides the first rigorous theoretical proof, supported by empirical validation, that RLHH can achieve the optimal expected runtime attainable by these two operators (up to lower-order terms) under appropriate parameter settings. This result refutes prior skepticism regarding RLHH’s capacity to learn effective heuristic selection policies. Furthermore, experimental results demonstrate that, at practical problem scales, RLHH outperforms the generalized randomized gradient hyper-heuristic, which also enjoys optimal theoretical runtime guarantees.
Existing learned heuristics often lack admissibility guarantees, rendering them unsuitable for optimal classical planning. This work proposes a large language model–driven evolutionary program synthesis framework that automatically generates interpretable programs tailored to individual planning domains. These programs construct pattern collections and, when combined with saturated cost partitioning, yield admissible, domain-specific heuristics. To the best of our knowledge, this is the first approach to produce learned, admissible, domain-dependent heuristics, preserving the optimality of A* while substantially reducing per-state evaluation overhead. Empirical results demonstrate that the method achieves coverage comparable to state-of-the-art domain-independent heuristics across multiple planning domains, offering both computational efficiency and strong theoretical guarantees.