🤖 AI Summary
This study addresses the reliance on expert knowledge in production scheduling modeling and the high code hallucination rates of general-purpose large language models (LLMs), which hinder practical deployment. To overcome these challenges, this work proposes an automated modeling framework integrating a multi-agent architecture with retrieval-augmented generation. Leveraging fine-tuning-free LLM agents, the method dynamically retrieves solver documentation via the Model Context Protocol to automatically translate natural language descriptions into executable constraint programming code, effectively bridging modeling and AI-driven decision-making. Experimental results demonstrate that the single-run success rate of generated scripts improves from 14.8% to 59.3%, reaching 80.6% for medium-complexity problems. These findings significantly mitigate hallucinations and validate the feasibility of LLM-assisted modeling.
📝 Abstract
Developing optimization models for production scheduling requires substantial expert effort. Research on large language models (LLMs) has followed two directions: specialized approaches for automated modeling, mostly for mixed-integer linear programming, which often rely on dedicated training or problem-specific architectures that limit industrial deployment; and agentic artificial intelligence for operational decision support, which generally assumes that the optimization model already exists. This study bridges both directions by assessing whether general-purpose LLMs, orchestrated as agents without task-specific training, can formulate and implement constraint programming models from natural-language problem descriptions. Singleagent and multi-agent architectures are integrated with a Model Context Protocol server that provides context-aware retrieval of solver documentation to mitigate hallucinations during implementation. Both are compared with a direct LLM baseline on six industry-oriented problems covering flow-shop, job-shop, flexible job-shop and resource-constrained warehouse scheduling, using three LLMs and assessing modeling accuracy, execution success, latency and token consumption. Formulation proves largely within reach of current LLMs, whereas implementation is the main barrier. The multi-agent workflow raises the share of scripts that run correctly as generated from 14.8% with a direct LLM call to 59.3%, reaching 80.6% on the four less complex problems, while tightly coupled intralogistics models remain an open challenge.