From Search to Signal: Online Post-Training in Automatic Heuristic Design

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that program evaluation signals inadequately guide model updates in LLM-driven automated heuristic design. To overcome this, we propose an online post-training framework that leverages Reinforcement Learning with Verifiable Rewards (RLVR) to translate program validity and performance into context-aware learning signals conditioned on prompts and evolutionary states. This approach mitigates the limitation of misjudging model capabilities based solely on end-to-end gains. Experimental results demonstrate that the proposed signal mapping strategy substantially enhances the heuristic generation capabilities of small-scale models. Furthermore, the benefits derived from online model updates surpass those achieved through additional search using frozen models.
📝 Abstract
Large language model (LLM)-based automatic heuristic design (AHD) iteratively proposes and refines heuristics, pairing design rationales with executable code. Task-specific evaluators assess programs; execution outcomes and performance scores guide search. Many AHD systems keep the generator frozen; EvoTune and Co-Evolution of Algorithms and Language Model (CALM) instead update it from evaluated candidates. When such outcomes drive reinforcement learning with verifiable rewards (RLVR), they create a search-coupled loop: the evaluated candidate stream supplies both search-state updates and training signals for the model that generates future candidates. Yet validity and performance do not uniquely determine useful model updates; converting them into learning signals must account for the prompt and evolving search state that produced each candidate. We formulate online post-training of small open-weight LLMs in AHD as context-dependent signal construction and develop alternative mappings from program validity, task performance, and generation context to update signals. Using shared evaluated rollouts and matched update budgets, controlled experiments across AHD tasks and model families compare these mappings with online post-training baselines, testing their effects on validity, performance among valid proposals, and the yield of valid proposals that improve under contextual comparisons. Complementary checkpoint, frozen-search, and live-system evaluations assess whether proposal-level gains appear in updated checkpoint behavior and subsequent search, rather than arising solely from accumulated search state. A resource-matched comparison under pre-specified cost accounting tests whether online updating adds value beyond additional search with a frozen generator. Together, this design avoids treating end-to-end search gains alone as evidence of stronger heuristic-design capabilities.
Problem

Research questions and friction points this paper is trying to address.

Automatic Heuristic Design
Online Post-Training
Reinforcement Learning with Verifiable Rewards
Signal Construction
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automatic Heuristic Design
Online Post-Training
Context-dependent Signal Construction
Reinforcement Learning with Verifiable Rewards
Controlled Evaluation
🔎 Similar Papers
No similar papers found.