Score
Designs and builds systems that convert problem descriptions—often in natural language—into formal optimization models by identifying decision variables, parameters, constraints, and objective functions and emitting solver-ready mathematical formulations or executable solver code. Analyzes and refines the mapping from input text or structured descriptions to model components to ensure correctness, feasibility, and compatibility with target solvers.
Mathematical optimization modeling heavily relies on domain experts and suffers from low automation. Method: This work systematically investigates how large language models (LLMs) can empower automated mathematical modeling, focusing on data synthesis, instruction fine-tuning, reasoning framework design, benchmark construction, and evaluation methodology. To address pervasive labeling errors (>40%) in mainstream benchmarks (e.g., OptiMath, MOBench), we conduct the first large-scale manual verification and cleaning, yielding the high-quality OptiClean dataset. Contribution/Results: Based on OptiClean, we establish the first fair, reproducible automated modeling leaderboard; release an open-source repository integrating datasets, code, literature, and an online evaluation platform; and provide a standardized evaluation framework, reliable benchmark, and scalable technical paradigm for LLM-driven modeling automation—significantly advancing the field’s standardization and rigor.
This work addresses the autoformulation problem—automatically translating natural-language problem descriptions into solvable mathematical optimization models. We propose the first LLM-driven Monte Carlo Tree Search (MCTS) framework for this task, enabling dynamic hypothesis generation and formal correctness evaluation. Our method integrates hierarchical optimization modeling representations, LLM-based semantic understanding, and MCTS-based search strategies. A key innovation is an equivalence-aware pruning mechanism that reduces search overhead by over 40%. Empirically, our approach achieves state-of-the-art performance on LP/MIP benchmarks, outperforming all existing baselines. LLM-assisted verification accelerates correctness assessment significantly. Moreover, this work formally defines the autoformulation task for the first time, establishing a scalable, automated paradigm to lower the barrier to optimization modeling and empower domain experts.
To address the accuracy bottleneck in automatic natural language-to-mixed-integer linear programming (NL-to-MILP) modeling—stemming from scarce high-quality annotated data and insufficient integration of domain expertise—this paper proposes an optimization-knowledge-enhanced large language model (LLM) framework. Our method comprises three core components: (1) a fine-grained, category-specific error analysis–driven data cleaning strategy; (2) a MILP-semantic-structured, class-aware multi-turn reasoning prompting framework; and (3) an iterative validation and refinement mechanism incorporating solver feedback. Extensive experiments across multiple foundational LLMs demonstrate an average 14.2-percentage-point improvement in modeling accuracy. Notably, robustness is significantly enhanced on critical subtasks—including complex constraint formulation and integer variable identification. The proposed approach establishes a new, interpretable, and solver-verified paradigm for AI-driven operations research modeling.
Many optimization problems in manufacturing, logistics, and healthcare remain reliant on manual heuristics due to the high modeling barrier for Mixed-Integer Linear Programming (MILP). Method: This paper proposes the first end-to-end MILP automation framework driven by natural language descriptions. It introduces a modular large language model (LLM) architecture integrating natural language understanding, program synthesis, code debugging, solution quality verification, and feedback-driven iterative refinement. Additionally, it establishes NLP4LP—the first long-horizon, complex LP benchmark dataset derived from natural language problem specifications. Contribution/Results: Experiments demonstrate that our framework achieves an accuracy gain of +12.3% over state-of-the-art methods on easy instances and +8.7% on hard instances—including those in NLP4LP—significantly advancing automated modeling and efficient solving of large-scale real-world optimization problems.
Small and medium-sized enterprises (SMEs) face high deployment costs, poor model reusability, and heavy reliance on expert domain knowledge when adopting combinatorial optimization decision-support systems. Method: This paper proposes the first fully automated, LLM-driven end-to-end paradigm that directly generates executable optimization code from natural language problem descriptions. Our approach integrates multi-stage prompt engineering, domain-knowledge injection, constraint-modeling guidance, and a cross-problem generalization evaluation framework—unifying problem understanding, mathematical modeling, solver integration, and code generation. Results: Evaluated on four canonical combinatorial optimization problem classes, the best-performing generator achieves over 70% syntactic correctness and 45% semantic functional correctness—substantially outperforming baseline methods. The core contribution lies in eliminating manual modeling bottlenecks, thereby enabling SMEs to construct optimization systems with minimal expertise, low entry barriers, and high model reusability.
This work addresses the challenge of automatically translating natural language descriptions of decision problems into executable optimization models by proposing an execution-aware modeling framework based on Autonomous Coding Agents (ACA). The approach ensures code executability through a sandboxed environment and introduces novel coordination mechanisms—including asymmetric verification loops, external memory reuse, minimum Bayes risk decoding, and self-consistency—to significantly enhance modeling robustness and accuracy. The system supports both interactive and fully autonomous operation, achieving state-of-the-art performance across nine standard optimization benchmarks and substantially outperforming existing methods on multiple datasets. These results validate the effectiveness of an architecture that treats ACA as a first-class abstraction for automated optimization modeling.
This study investigates whether large language models (LLMs) underperform in generating domain-specific languages—such as AMPL for algebraic modeling—compared to general-purpose programming languages like Python, specifically within mathematical optimization contexts. To address this, the authors propose EXEOS, a method that leverages LLMs to translate natural language descriptions into either AMPL or Python code, augmented with a solver-feedback-driven iterative refinement mechanism to enhance executability and correctness. The first systematic comparison of its kind demonstrates that, across public benchmarks and real-world Kinaxis supply chain cases, LLM-generated AMPL code matches or even surpasses Python in quality. These findings affirm the competitiveness of domain-specific languages in specialized optimization tasks and highlight the critical role of solver-in-the-loop iterative refinement in improving the generation of formal specifications.
This work addresses the lack of evaluation benchmarks aligned with real-world industrial-scale optimization models for natural language to optimization modeling approaches, which hinders reliable performance assessment on large-scale practical problems. To bridge this gap, the authors propose a structure-aware inverse construction method that recovers compact model structures from real mixed-integer linear programming (MILP) instances in MIPLIB 2017 and generates semantically precise natural language descriptions. This yields MIPLIB-NL, the first industrial-grade benchmark supporting model-data separation, comprising 223 one-to-one reconstructed instances. Through expert review and human-in-the-loop iterative validation, the benchmark reveals significant performance degradation of current systems on authentic industrial-scale problems and uncovers failure modes invisible to toy-scale benchmarks, thereby establishing a reliable standard for evaluating automated optimization modeling.
This work proposes LLMize, a framework for complex numerical optimization problems where formalizing constraints and heuristics is challenging. By treating the optimization process as a black box, LLMize leverages large language models (LLMs) to generate candidate solutions in natural language, iteratively refining them through evaluations from an external objective function and feedback. Notably, it enables direct injection of constraints and domain knowledge via natural language—bypassing the need for mathematical programming or metaheuristic design—and thereby substantially lowers the barrier to tackling intricate optimization tasks. The framework integrates iterative prompting, in-context learning, OPRO, and a hybrid strategy inspired by evolutionary algorithms and simulated annealing. Empirical validation across diverse tasks—including convex optimization, linear programming, TSP, hyperparameter tuning, and nuclear fuel assembly layout—demonstrates its efficacy: while underperforming classical solvers on simple problems, it exhibits unique practical value in complex domains.