🤖 AI Summary
This study addresses the challenge that existing Mixed-Integer Linear Programming (MILP) instance generation methods struggle to simultaneously ensure feasibility and computational hardness while lacking explicit hardness metrics. To overcome these limitations, this work proposes a reinforcement learning framework driven by solver feedback. Specifically, it constructs reward signals based on branch-and-bound node counts and optimality gaps, employing the Group Relative Policy Optimization (GRPO) algorithm to fine-tune large language models such as Gemma and Qwen. Furthermore, an asymmetric seedless self-play mechanism is introduced to iteratively guide the model in generating feasible yet computationally challenging instances. Experimental results demonstrate that the proposed approach significantly increases SCIP search node counts and post-cut gaps, effectively expanding the difficulty spectrum of generated instances while improving their feasibility rates.
📝 Abstract
Generating optimization instances that are both feasible and computationally challenging is crucial for benchmarking solvers and training learning-based optimization algorithms. Existing non-LLM generators rely on seed instances or parameter tuning, resulting in high test-time computational cost, while existing LLM generators lack explicit hardness measures. Recent reinforcement learning methods with verifier feedback evaluate only binary correctness, which is misaligned with generating challenging problems. We note that an optimization solver reports the cost of solving at several stages of its pipeline, and leverage this to design a reward that scores both the solvability and the hardness of generated problems, measured by branch-and-bound nodes and post-cut relaxation gaps. Our key idea is a challenger-solver asymmetric self-play approach, where an LLM challenger generates progressively harder instances and the solver verifies feasibility and hardness, so no seed or training MILP instances are required. We fine-tune Gemma-4-12B and Qwen3.5-4B with GRPO and a size curriculum into OptiScribe-12B and OptiScribe-4B, which generate feasible yet challenging MILP problems from natural language instructions. On capacitated facility location and max-cut, OptiScribe-12B raises median SCIP search nodes by 1.7-5x and post-cut gaps by 1.1-1.7x over its base model and improves the feasibility rate on facility location by 9-19 points, while OptiScribe-4B raises median nodes by up to 15.6x. The problems cover a wider difficulty range than public benchmarks of the same size, follow instructions on density and difficulty, and can tune solver settings for families that public libraries lack. These results indicate that optimization-specific rewards, used in self-play mode, can teach LLMs to generate high-difficulty optimization benchmarks. We will release our code and models publicly on acceptance.