AI Papers Reader

Personalized digests of latest AI research

View on GitHub

Verification Gate Helps AI Agents Write PLC Code That Runs, Researchers Report

Programmable logic controllers (PLCs), the industrial computers that run factory lines, power plants and water-treatment facilities, are usually programmed in IEC 61131-3 languages, including the text-based Structured Text (ST). Large language models can already write standalone program units for them. Production control logic, however, must fit into an existing project, reuse its modules, and behave correctly once the plant is running. A program can compile and still misconfigure a timer or miss a reset. According to the authors, few studies have measured this runtime behavior across methods and models.

The researchers present SemaPLC, an agent harness, meaning a wrapper that lets a language model plan, edit and call tools in a loop. Its central rule governs when the agent may stop. Instead of accepting its own judgment that the output is adequate, SemaPLC declares a task complete only when logged external checks confirm it. These checks cover a clause-by-clause audit against the requirement, compilation, and a live runtime test in which the program is deployed to a controller, fed scenario inputs, and its outputs sampled. Any edit voids earlier results, and each failed check receives at most two repair rounds.

The paper gives one case that illustrates the failure mode. A requirement states that low flow (below 50 kg/hr) should set a setpoint to 500, and that a transmitter fault should set it to 2500. A baseline, Agents4PLC, produced code that compiled, but its low-flow assignment overwrote the fault default. The authors inferred this from the source logic rather than from a runtime trace. SemaPLC’s first candidate also compiled and had a similar priority error. When the runtime test forced the flow input to 30 kg/hr, the output was 2500 rather than the expected 500. The repair selected the output according to cause, and re-verification confirmed both abnormal cases.

On a 117-task function-level benchmark, SemaPLC had the highest strict verified pass rate on all seven backbone models, averaging 72.6%, compared with 63.9% for the strongest baseline, Agents4PLC. That is 8.8 percentage points higher. Running the same models without the harness averaged 55.3%, so the harness added 17.3 points on average.

On a 65-task project track, the largest gap appeared at runtime. SemaPLC’s mean integrated compilation rate was 89.4%, against 58.7 to 81.5 for the baselines. Static scores, which check the program text against assertions, were close: 81.6 for SemaPLC against 71.7 to 75.7. Dynamic scores, which compare executed traces with a reference, averaged 52.2, compared with 22.4 to 31.4 for the baselines. SemaPLC led on dynamic behavior for all seven models, but on GPT-5.5 its lead over Agents4PLC narrowed to 1.8 points (65.4 against 63.6).

The approach has costs. On the project track, SemaPLC averaged 34.1 model requests per task, against 6.9 for Agents4PLC, although wall-clock times were similar (347 seconds against 344). The dynamic tests use at most six scenarios from a hidden reference, so behavior under other conditions remains unmeasured. The authors also report that none of the 174 properties in timer-bearing programs received a conclusive result from their formal model checker, which is one reason they rely on runtime testing for those cases.

The work suggests that static scores can mask large differences in runtime behavior, and that industrial code benchmarks should execute generated logic. The results cover specific tasks, models and test scenarios, and the code is open-sourced, so independent replication is possible.