AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise

📅 2025-09-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Enterprise-grade multi-agent systems lack empirical studies on cross-architectural interactions; existing work typically evaluates components in isolation, obscuring the true efficacy of configuration combinations. Method: We introduce the first enterprise-oriented agent architecture benchmark, systematically evaluating 18 configurations across four core dimensions—orchestration strategy, prompt engineering, memory architecture, and tool integration—using ReAct versus function-calling paradigms, multi-dimensional ablation analysis, and realistic task frameworks. Contribution/Results: We identify strong architectural preferences, challenging the “one-size-fits-all” design paradigm; the best-performing configuration achieves only 35.3% and 70.8% success rates on complex and simple tasks, respectively, revealing fundamental performance bottlenecks. Our findings advance the development of enterprise-tailored agent architectures grounded in empirical evidence and systematic evaluation.

Technology Category

Cognitive Modeling & Cognitive Systems: Agent ArchitecturesMultiagent Systems: Agent/AI Theories and ArchitecturesSearch and Optimization: Algorithm Configuration

Application Category

Search and Retrieval-Augmented AI: Agentic searchSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: Architectures and workflows that use LLMs for crowd work
📝 Abstract
While individual components of agentic architectures have been studied in isolation, there remains limited empirical understanding of how different design dimensions interact within complex multi-agent systems. This study aims to address these gaps by providing a comprehensive enterprise-specific benchmark evaluating 18 distinct agentic configurations across state-of-the-art large language models. We examine four critical agentic system dimensions: orchestration strategy, agent prompt implementation (ReAct versus function calling), memory architecture, and thinking tool integration. Our benchmark reveals significant model-specific architectural preferences that challenge the prevalent one-size-fits-all paradigm in agentic AI systems. It also reveals significant weaknesses in overall agentic performance on enterprise tasks with the highest scoring models achieving a maximum of only 35.3% success on the more complex task and 70.8% on the simpler task. We hope these findings inform the design of future agentic systems by enabling more empirically backed decisions regarding architectural components and model selection.
Problem

Research questions and friction points this paper is trying to address.

Evaluates agent architectures in enterprise multi-agent systems
Examines orchestration strategy and memory architecture interactions
Benchmarks agent performance on complex enterprise tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Benchmark evaluates 18 agent configurations
Examines four critical agentic system dimensions
Reveals model-specific architectural preferences empirically
T
Tara Bogavelli
ServiceNow
R
Roshnee Sharma
ServiceNow
H
Hari Subramani
ServiceNow