Evaluation and Benchmarking of LLM Agents: A Survey

📅 2025-07-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current LLM agent evaluation lacks a systematic framework, particularly neglecting enterprise-specific requirements such as role-based access control, regulatory compliance, and long-horizon interactive behavior. To address this gap, we conduct a comprehensive literature review and propose the first two-dimensional evaluation taxonomy: one axis captures *objective dimensions*—including behavior, capability, reliability, and security—while the other captures *process dimensions*—encompassing interaction paradigms, benchmark datasets, evaluation metrics, and toolchains. Crucially, our taxonomy explicitly incorporates enterprise challenges, exposing critical shortcomings in existing work regarding holistic coverage, scalability, and real-world applicability. The framework provides researchers and practitioners with a structured, actionable reference for designing, evaluating, and deploying LLM agents in complex, mission-critical environments. It advances the field toward trustworthy, production-ready LLM agent systems. (138 words)

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Agent ArchitecturesMultiagent Systems: Modeling other Agents

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
The rise of LLM-based agents has opened new frontiers in AI applications, yet evaluating these agents remains a complex and underdeveloped area. This survey provides an in-depth overview of the emerging field of LLM agent evaluation, introducing a two-dimensional taxonomy that organizes existing work along (1) evaluation objectives -- what to evaluate, such as agent behavior, capabilities, reliability, and safety -- and (2) evaluation process -- how to evaluate, including interaction modes, datasets and benchmarks, metric computation methods, and tooling. In addition to taxonomy, we highlight enterprise-specific challenges, such as role-based access to data, the need for reliability guarantees, dynamic and long-horizon interactions, and compliance, which are often overlooked in current research. We also identify future research directions, including holistic, more realistic, and scalable evaluation. This work aims to bring clarity to the fragmented landscape of agent evaluation and provide a framework for systematic assessment, enabling researchers and practitioners to evaluate LLM agents for real-world deployment.
Problem

Research questions and friction points this paper is trying to address.

Evaluating LLM agents' behavior, capabilities, reliability, and safety
Addressing enterprise challenges like data access and compliance
Developing holistic and scalable evaluation methods for real-world use
Innovation

Methods, ideas, or system contributions that make the work stand out.

Two-dimensional taxonomy for LLM agent evaluation
Addresses enterprise-specific evaluation challenges
Proposes future holistic scalable evaluation methods
🔎 Similar Papers
No similar papers found.
SAP Labs | SAP Labs
M
Mahmoud Mohammadi
SAP Labs, Bellevue, WA, USA
Y
Yipeng Li
SAP Labs, Bellevue, WA, USA
J
Jane Lo
SAP Labs, Palo Alto, CA, USA
W
Wendy Yip
SAP Labs, Palo Alto, CA, USA