ANI-Gamut: Benchmarking Agent Reliability across the Gamut of Agent-Network Interface Abstractions

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过ANI-Gamut平台在不同抽象层级的代理-网络接口上测试代理可靠性,使用实验方法比较了从原始CLI到类型化事务服务意图等多个层次的效果。
📝 Abstract
Large Language Model (LLM) agents are increasingly trusted to operate live networks: they read state, change configuration, and verify the result. A first-order question is left implicit: at which level of abstraction should the agent operate? We make the interface-abstraction level an explicit, controlled experimental variable, organizing agent-network interfaces into a spectrum from raw CLI (A0) through bounded wrappers (A1) and standardized model-driven configuration (A2) to typed transactional service intent (A3-T), reconciled source-of-truth automation (A3-R), and their combination (A4). We present ANI-Gamut, a reproducible, open-source playground that exposes the same task at several levels on a single, densely populated brownfield substrate, where many coexisting services share resources and collateral damage actually arises. We instantiate and measure four points of the spectrum and describe the others only at the conceptual level, and record three dependent variables as the substrate is stressed by injected faults: task reliability, collateral damage against pre-existing tenants, and operational cost. In a pilot with small run counts (n=20, n=6 and n=4 per level), the interface level moves reliability and cost sharply: on a live change-set task a raw-shell agent fails on all six seeds while a typed transactional interface succeeds on all six (paired McNemar p=0.031, on six discordant pairs), at roughly an order of magnitude less cost, as the engineering effort migrates from the agent to a reusable transaction layer. Collateral damage, by contrast, is absent at every level, whether benign or under faults: in a tenant-isolated substrate the agents fail safe, and the blast radius is held by the substrate's isolation, which moves the safety question from the agent to the substrate.
Problem

Research questions and friction points this paper is trying to address.

Agent-Network Interface Abstractions
Large Language Model (LLM) agents
task reliability
collateral damage
operational cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

ANI-Gamut
Agent-Network Interface Abstractions
Large Language Model Agents
Task Reliability
Collateral Damage
🔎 Similar Papers
No similar papers found.
L
Lorenzo Bracciale
University of Rome Tor Vergata, Italy
P
Pierpaolo Loreti
University of Rome Tor Vergata, Italy
A
Andrea Mayer
University of Rome Tor Vergata, Italy
Stefano Salsano
Stefano Salsano
Dip. Ingegneria Elettronica - University of Rome Tor Vergata
computer networksSDNmobile computingsoftware defined networkingnetwork function virtualization
W
Wim Henderickx
Nokia, Belgium