SemOPT: Fixing Semantic Errors in LLM-based Optimization Modeling via Reward-Guided Search

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insidious semantic errors in code generated by large language models for optimization modeling, where outputs are syntactically executable yet violate the underlying problem intent. To tackle this challenge, we propose SemOPT, a novel framework that introduces the first semantic reward mechanism designed to distinguish faithful solutions from deceptively correct ones. Furthermore, SemOPT incorporates an adaptive hierarchical reward-guided search strategy to enable the automatic detection and correction of semantic deviations. Extensive experiments demonstrate that SemOPT achieves state-of-the-art performance across seven benchmarks, yielding an average accuracy improvement of 7.6% on complex datasets.
📝 Abstract
Operations research supports decision-making in domains such as energy, economics, and healthcare. Solving operations research problems typically begins with optimization modeling, which translates a natural-language problem description into executable solver code. LLMs offer a promising way to automate this process, but they remain prone to errors. In practice, these errors can be divided into two categories: syntactic errors refer to solver code that fails to run successfully or is judged infeasible by the solver; semantic errors refer to solver code that successfully returns an objective value but violates the intent of the original problem. Since semantic errors do not trigger runtime failures, they are difficult to detect and rectify. To address this problem, we introduce SemOPT, a semantic-guided framework for correcting LLM-based optimization models. SemOPT combines a semantic reward model that distinguishes faithful math models from plausible but incorrect ones with an adaptive correction system that applies hierarchical reward-guided search over the modeling space. Experiments on seven optimization modeling benchmarks show that SemOPT establishes a new state of the art and achieves an average 7.6% accuracy improvement over the strongest baseline on complex datasets.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Optimization Modeling
Semantic Errors
Operations Research
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Error Correction
Reward-Guided Search
Optimization Modeling
Large Language Models
Semantic Reward Model