PLCWorld: Benchmarking LLM-Generated PLC Programs in Closed-Loop Plant Simulation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of benchmarks for evaluating the safety and functionality of PLC programs generated by large language models (LLMs) by constructing a closed-loop evaluation environment that couples structured text execution with plant simulation. We introduce a novel difficulty grading system based on control dependency scope, alongside one hundred synthetic tasks and automated protocols enabling independent quantification of task success rates and safety violation rates. Experiments reveal that GPT-5.5 achieves only a 25% success rate on high-difficulty tasks, exposing significant deficiencies in current models regarding safety compliance. By open-sourcing the complete environment and benchmark data, this work bridges a critical gap in industrial-grade closed-loop verification of LLM-generated code.
📝 Abstract
Programmable logic controllers (PLCs) coordinate industrial equipment by reading sensor inputs and issuing control commands. Evaluating whether large language model (LLM)-generated PLC programs satisfy task requirements and safety constraints requires observing how their commands affect device and workpiece states. We introduce PLCWorld, a common closed-loop execution environment and benchmark that couples Structured Text (ST) execution with simulated plant responses and sensor feedback. Grounded in control relations identified in industrial PLC programs and engineering documentation, PLCWorld contains 100 synthetic tasks and 473 registered task-condition pairs across Motion Control and Material Handling, with difficulty defined by control-dependency scope. A common protocol reports Task Success and Safety Violation separately. Validation combines practitioner review, reference and alternative programs, targeted counterexamples, specification-evaluator alignment checks, and comparisons with independent ST runtimes. Reference and alternative programs satisfy their applicable cases, while all 542 targeted counterexamples activate their designated evaluator rules under at least one registered condition. Execution Gap relates submission-profile acceptance to subsequent task failure or observed Safety Violation. Across the constructed task groups, direct GPT-5.5 achieves 82.70% Task Success on Easy cases but 25.10% on Hard cases. Evaluations of six LLMs and four adapted generation-and-verification workflows further expose differences between completion, safety, and generation cost. Our code, simulation environment, benchmark tasks, and baseline implementations are publicly available at https://yunji0516.github.io/PLCWorld/.
Problem

Research questions and friction points this paper is trying to address.

Programmable Logic Controllers
Large Language Models
Closed-Loop Simulation
Benchmark
Safety Violation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Closed-loop simulation
PLC benchmark
Structured Text
Safety evaluation
Large language models