LLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale Optimization

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that large language models (LLMs) face in handling optimization tasks of industrial scale and structural diversity. To this end, it proposes a Strategy-Diverse Reinforcement Learning (SDRL) framework that trains open-source LLMs into adaptive meta-solvers. Methodologically, the framework introduces a correctness-gated hierarchical diversity reward mechanism to mitigate policy collapse, alongside a mixed-format training paradigm that ensures compatibility with both textual and file-based input instances. Experimental results demonstrate that SDRL outperforms state-of-the-art closed-source models, including DeepSeek-V4-Pro and GPT-5.5, across standard benchmarks and industrial-scale optimization tasks. These findings substantiate the potential of open-source models for solving complex optimization problems.
📝 Abstract
Scaling LLM-based optimization from textbook-scale instances to real-world, industrial tasks remains a critical open challenge. Existing approaches are predominantly evaluated on small, self-contained textual problems and often commit to a solver-integrated paradigm, limiting their ability to handle the scale and structural diversity of practical optimization workloads. In this work, we propose a practical framework for training open-source LLMs to tackle real-world, industrial-scale optimization. We first show empirically that solver-integrated reasoning, exact combinatorial algorithm, and heuristic search exhibit complementary strengths across different problem structures and scales. Motivated by this, we introduce Strategy-Diverse Reinforcement Learning (SDRL), which trains LLMs as adaptive optimization meta-solvers. SDRL leverages this complementarity through a correctness-gated hierarchical diversity reward that promotes robust exploration across varying strategies and within each strategy, effectively preventing premature strategy collapse. We further introduce a mixed-format training scheme that jointly supports both self-contained textual problems and file-grounded instances. Across comprehensive evaluations, our framework outperforms existing fine-tuned methods and frontier models including DeepSeek-V4-Pro and GPT-5.5, both on average across benchmarks and on industrial-scale optimization tasks.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Industrial-scale Optimization
Combinatorial Optimization
Meta-Solvers
Innovation

Methods, ideas, or system contributions that make the work stand out.

Strategy-Diverse Reinforcement Learning
Adaptive Meta-Solvers
Hierarchical Diversity Reward
Mixed-Format Training
Industrial-Scale Optimization
🔎 Similar Papers
No similar papers found.