🤖 LLM Agents vs. Factory PLCs: Success in Only a Third of Attempts
A team from Chongqing University and Zhejiang University published arXiv preprint 2608.26882 on August 27 — the first benchmark of autonomous LLM agents on real PLCs in hardware-in-the-loop mode: 4 controllers, 4 loads, 5 LLM families. Only 75 out of 240 attempts (31.3%) achieved a sustained physical effect.
🌍 The benchmark localizes the failure stage: reading → writing → maintaining the effect. Process observability is a risk lever: success after writing increases from 44.2% to 64.0%. For PLC vendors, these are points for evaluating defenses: telemetry restriction, detection of vendor-native writing.
👤 Today's agents against factory hardware most often fail before the first legitimate reading operation. The code and software pipeline for reproduction are open — the benchmark can be reproduced without an industrial stand.
Source 1: https://arxiv.org/abs/2608.26882 Source 2: https://arxiv.org/pdf/2608.26882
