🤖 AI Summary
Domain-agnostic iterative dynamic programming (DIDP) for combinatorial optimization suffers from low search efficiency due to its reliance on manually designed dual bounds. Method: This paper proposes a reinforcement learning–based framework for automatically constructing general-purpose heuristics. It is the first to integrate Deep Q-Networks (DQN) and Proximal Policy Optimization (PPO) into the DIDP paradigm, aligning the RL state-transition dynamics with the Bellman equation structure to enable data-driven heuristic learning without expert knowledge. Contribution/Results: Experiments across four benchmark domains and three task types demonstrate that the proposed method significantly reduces runtime compared to standard DIDP, while consistently outperforming problem-specific greedy heuristics under identical node expansion budgets. These results validate both the method’s strong generalization capability and its superior search efficiency.
📝 Abstract
Domain-Independent Dynamic Programming (DIDP) is a state-space search paradigm based on dynamic programming for combinatorial optimization. In its current implementation, DIDP guides the search using user-defined dual bounds. Reinforcement learning (RL) is increasingly being applied to combinatorial optimization problems and shares several key structures with DP, being represented by the Bellman equation and state-based transition systems. We propose using reinforcement learning to obtain a heuristic function to guide the search in DIDP. We develop two RL-based guidance approaches: value-based guidance using Deep Q-Networks and policy-based guidance using Proximal Policy Optimization. Our experiments indicate that RL-based guidance significantly outperforms standard DIDP and problem-specific greedy heuristics with the same number of node expansions. Further, despite longer node evaluation times, RL guidance achieves better run-time performance than standard DIDP on three of four benchmark domains.