🤖 AI Summary
This study addresses the critical limitation that optimization models generated by large language models (LLMs), while often correct, suffer from low solving efficiency, thereby hindering their practical deployment. To tackle this issue, the work introduces the concept of an "efficiency gap" and proposes OptTips, a knowledge base integrating expert heuristics, alongside OptDachshund, a multi-agent framework. Furthermore, it establishes EfficientOpt, a paired benchmark designed to jointly evaluate correctness and computational efficiency. Through a systematic evaluation of eleven LLMs, the study reveals a significant efficiency gap, demonstrating that 57% of functionally correct programs are solved more slowly than expert implementations. These findings underscore the necessity for LLM-based modeling approaches to simultaneously prioritize both correctness and computational efficiency to ensure real-world applicability.
📝 Abstract
Optimization modeling formulates real-world decision problems as mathematical programs that solvers can use to find optimal decisions. Large language models (LLMs) can automate this process, but the resulting correct formulations can require substantial time and memory to construct and solve, limiting practical scalability. Therefore, we systematically investigate whether LLMs can identify problem structure from natural-language descriptions and apply suitable optimization modeling techniques to generate mathematical models and solver code that solve the problems correctly and efficiently. To this end, we first curate OptTips, a knowledge base of 50 expert modeling techniques in eight families. Using this knowledge, we develop OptDachshund, a multi-agent framework that transforms problems from existing optimization benchmarks into new tasks for evaluating LLMs' use of modeling techniques. It constructs conventional and expert mathematical models with solver code for the same task and data, providing baselines for correctness and computational cost. The resulting EfficientOpt benchmark contains 561 expert-reviewed tasks with paired reference implementations. Evaluation of 11 representative LLMs reveals an efficiency gap on correctly solved tasks with comparable measurements: for every LLM, most generated programs take longer to solve than their expert counterparts. Within the comparable reference-size subset, 57\% of programs with correct objective values and fewer variables and linear constraints have longer recorded solver times. Case studies show that different modeling techniques can achieve the same optimal value at similar recorded cost. Faster solving may not reduce execution time if the code takes longer to prepare data and build the model. LLM optimization modeling should therefore be evaluated for both correctness and computational efficiency.