Reinforcement Learning-based Heuristics to Guide Domain-Independent Dynamic Programming

📅 2025-03-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Domain-agnostic iterative dynamic programming (DIDP) for combinatorial optimization suffers from low search efficiency due to its reliance on manually designed dual bounds. Method: This paper proposes a reinforcement learning–based framework for automatically constructing general-purpose heuristics. It is the first to integrate Deep Q-Networks (DQN) and Proximal Policy Optimization (PPO) into the DIDP paradigm, aligning the RL state-transition dynamics with the Bellman equation structure to enable data-driven heuristic learning without expert knowledge. Contribution/Results: Experiments across four benchmark domains and three task types demonstrate that the proposed method significantly reduces runtime compared to standard DIDP, while consistently outperforming problem-specific greedy heuristics under identical node expansion budgets. These results validate both the method’s strong generalization capability and its superior search efficiency.

Technology Category

Search and Optimization: Heuristic SearchPlanning, Routing, and Scheduling: Learning for Planning and SchedulingConstraint Satisfaction and Optimization: Constraint Learning and Acquisition

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Domain-Independent Dynamic Programming (DIDP) is a state-space search paradigm based on dynamic programming for combinatorial optimization. In its current implementation, DIDP guides the search using user-defined dual bounds. Reinforcement learning (RL) is increasingly being applied to combinatorial optimization problems and shares several key structures with DP, being represented by the Bellman equation and state-based transition systems. We propose using reinforcement learning to obtain a heuristic function to guide the search in DIDP. We develop two RL-based guidance approaches: value-based guidance using Deep Q-Networks and policy-based guidance using Proximal Policy Optimization. Our experiments indicate that RL-based guidance significantly outperforms standard DIDP and problem-specific greedy heuristics with the same number of node expansions. Further, despite longer node evaluation times, RL guidance achieves better run-time performance than standard DIDP on three of four benchmark domains.
Problem

Research questions and friction points this paper is trying to address.

Enhance Domain-Independent Dynamic Programming with RL heuristics.
Develop RL-based guidance using Deep Q-Networks and Proximal Policy Optimization.
Improve search efficiency and runtime performance in combinatorial optimization.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning guides Dynamic Programming
Deep Q-Networks for value-based guidance
Proximal Policy Optimization for policy-based guidance
M
Minori Narita
University of Toronto, 5 King’s College Road, Toronto, Ontario, Canada
R
Ryo Kuroiwa
National Institute of Informatics, 2-1-2 Hitotsubashi, Chiyoda-ku, Tokyo, Japan
J. Christopher Beck
J. Christopher Beck
Professor of Mechanical and Industrial Engineering, University of Toronto
Constraint ProgrammingAutomated PlanningHeuristic SearchArtificial IntelligenceOperations Research