🤖 AI Summary
This study addresses the limitations of dynamic embeddings in neural combinatorial optimization, where representations require frequent reconstruction and reliance on external labels or search space pruning. We propose a purely reinforcement learning-based dynamic embedding framework that achieves selective memory propagation through multi-step computation. By integrating attention-based pre-fusion with post-adaptive gated updates, the method efficiently consolidates historical memory with current states, facilitating shallow policy network learning without external supervision or pruning operations. Experimental results demonstrate that this framework generates high-quality solutions across four classes of combinatorial optimization problems, scaling effectively to large instances comprising millions to tens of millions of nodes while exhibiting superior generalization capability and computational efficiency.
📝 Abstract
Constructive neural combinatorial optimization (NCO) has emerged as a promising paradigm that learns to construct solutions to combinatorial optimization problems (COPs) step by step, which reduces reliance on handcrafted rules and enables fast inference. While many methods with dynamic embeddings generalize well, they typically rebuild subproblem representations from scratch at each step using deep attention stacks. Many high-performing methods in this category rely on solution labels or pseudo-labels for efficient training, or on aggressive search space pruning during reinforcement learning (RL). To address these limitations, we propose Memory-in-the-Loop (MiLoop), a purely RL-based constructive framework that leverages the multi-step computation already required by a rollout for selective memory propagation. Each rollout provides solution-quality feedback for learning while propagating historical representations, thereby enabling a shallow policy to learn effective dynamic embeddings without external solution labels or training-time search-space pruning. Specifically, MiLoop fuses current embeddings with historical memory before the attention layers and applies adaptive gated updates afterward. The updated representations support both current decisions and stepwise reuse. Extensive experiments across four COPs demonstrate that MiLoop consistently produces high-quality solutions on instances ranging from 100 to 10 million nodes, highlighting its strong generalization ability.