🤖 AI Summary
Enterprise-grade multi-agent systems lack empirical studies on cross-architectural interactions; existing work typically evaluates components in isolation, obscuring the true efficacy of configuration combinations. Method: We introduce the first enterprise-oriented agent architecture benchmark, systematically evaluating 18 configurations across four core dimensions—orchestration strategy, prompt engineering, memory architecture, and tool integration—using ReAct versus function-calling paradigms, multi-dimensional ablation analysis, and realistic task frameworks. Contribution/Results: We identify strong architectural preferences, challenging the “one-size-fits-all” design paradigm; the best-performing configuration achieves only 35.3% and 70.8% success rates on complex and simple tasks, respectively, revealing fundamental performance bottlenecks. Our findings advance the development of enterprise-tailored agent architectures grounded in empirical evidence and systematic evaluation.
📝 Abstract
While individual components of agentic architectures have been studied in isolation, there remains limited empirical understanding of how different design dimensions interact within complex multi-agent systems. This study aims to address these gaps by providing a comprehensive enterprise-specific benchmark evaluating 18 distinct agentic configurations across state-of-the-art large language models. We examine four critical agentic system dimensions: orchestration strategy, agent prompt implementation (ReAct versus function calling), memory architecture, and thinking tool integration. Our benchmark reveals significant model-specific architectural preferences that challenge the prevalent one-size-fits-all paradigm in agentic AI systems. It also reveals significant weaknesses in overall agentic performance on enterprise tasks with the highest scoring models achieving a maximum of only 35.3% success on the more complex task and 70.8% on the simpler task. We hope these findings inform the design of future agentic systems by enabling more empirically backed decisions regarding architectural components and model selection.