π€ AI Summary
This study addresses the error-proneness of local small models and privacy leakage risks associated with cloud API calls in network automation. To mitigate these issues, it proposes a βverifiabilityβ criterion and the Touchstone pipeline, which integrates seven small language models (1β8B parameters) with task-specific intrinsic verification mechanisms and a tiered reasoning architecture. Locally generated outputs are filtered through deterministic checks, and only those failing verification are escalated to frontier large language models. Experimental results demonstrate that this framework achieves 93.8%β98.6% accuracy on tasks such as conflict detection, while requiring escalation for merely ~16% of inputs. Consequently, the proposed approach substantially reduces reliance on cloud services, effectively balancing privacy preservation with high predictive accuracy.
π Abstract
Sending every network-automation input to a third-party frontier LLM exports sensitive artifacts such as production configurations, topologies, and logs. Querying small language models (SLMs) locally avoids this egress, but SLM outputs can be error-prone for direct use. This work introduces checkability as a criterion for determining which tasks are suitable for local inference. A task is checkable when it exposes a cheap, deterministic test - an intrinsic check - that rejects outputs violating a necessary correctness condition. We instantiate this idea in Touchstone, a local-first pipeline that uses seven off-the-shelf SLMs (1-8B parameters) to generate candidates, uses task-specific intrinsic checks to reject responses, and escalates unresolved inputs to a frontier LLM. On conflict detection and intent translation tasks, Touchstone reaches 98.6% and 93.8% end-to-end accuracy while escalating only 16% and 17% of inputs, respectively. On TeleQnA, a knowledge-only control that has no task-specific intrinsic checks, Touchstone is unable to match the accuracy of the frontier baseline. Our results support a simple deployment rule: keep inference local when task semantics support precise, low-cost checks; escalate the rest.