🤖 AI Summary
This study addresses the lack of rapid, label-free attribution methods for attacks, faults, and anomalies in water supply SCADA alarms. To this end, it constructs four EPANET-based attribution benchmarks and proposes Jev, a training-free probabilistic model enabling second-level preliminary screening, which is integrated with rule trees and large language models (LLMs) into a gated cascaded review architecture. The proposed approach effectively resolves supervised model failures under few-shot conditions. It achieves in-distribution performance comparable to rule trees while surpassing supervised classifiers on unseen events. Furthermore, the framework accelerates decision-making by 20–40 times relative to standalone LLMs and reduces LLM invocations by 35%–38% without compromising accuracy.
📝 Abstract
When a SCADA alarm is raised in a water distribution network, operators must decide quickly whether it reflects a cyberattack, a physical fault, a normal transient or a faulty sensor. Supervised classifiers need labelled incidents that utilities rarely have, and frontier large language models (LLMs) take tens of seconds per decision. We tested whether Jev, a training-free model that returns class probabilities in about one second, can serve as the first tier of this triage. On a four-class cause-attribution benchmark built on the C-Town network in EPANET, Jev was compared with a hand-written rule tree, a supervised classifier and seven cloud LLMs on identical evidence in four sealed, pre-registered rounds. With only a label-free prior correction, Jev matched the rule tree (macro-F1 0.62-0.64 against 0.56-0.61 in distribution) and exceeded the supervised classifier by 0.36-0.42 on event subtypes absent from its labels, in all four rounds, and it outperformed the classifier whenever fewer than about four labelled events per class were available. Jev also decided 20-40 times faster than frontier LLMs. Accepting only benign Jev verdicts confirmed by the rule tree spared an LLM reviewer 35-38% of windows on fresh sealed sets without loss of macro-F1. Transferred unchanged to two further networks, this gated cascade stayed within the non-inferiority margin of its reviewer on all four sets. A fast, training-free screen can therefore take over about a third of the review load in SCADA anomaly triage while preserving the accuracy of deliberate review.