🤖 AI Summary
This work addresses the lack of systematic empirical comparison among tool integration and agent delegation protocols in multi-agent systems, which hinders informed architectural choices for complex task orchestration. We propose the first evaluation framework specifically designed for assessing multi-agent communication protocols in task orchestration, establishing a standardized benchmark to compare pure tool-integration, pure multi-agent delegation, and hybrid architectures across three levels of task complexity. Leveraging a standardized query set, end-to-end metric collection, and real-world deployment environments, we quantitatively analyze performance along multiple dimensions—including response time, context consumption, cost, error recovery, and implementation complexity. Our experiments reveal critical trade-offs among the protocols in terms of efficiency, resource overhead, and robustness, offering data-driven guidance for practical system design.
📝 Abstract
Context. Nowadays, artificial intelligence agent systems are transforming from single-tool interactions to complex multi-agent orchestrations. As a result, two competing communication protocols have emerged: a tool integration protocol that standardizes how agents invoke external tools, and an inter-agent delegation protocol that enables autonomous agents to discover and delegate tasks to one another. Despite widespread industry adoption by dozens of enterprise partners, no empirical comparison of these protocols exists in the literature. Objective. The goal of this work is to develop the first systematic benchmark comparing tool-integration-only, multi-agent delegation, and hybrid architectures across standardized queries at three complexity levels, and to quantify the trade-offs in response time, context window consumption, monetary cost, error recovery, and implementation complexity.