🤖 AI Summary
Current LLM agent evaluation lacks a systematic framework, particularly neglecting enterprise-specific requirements such as role-based access control, regulatory compliance, and long-horizon interactive behavior. To address this gap, we conduct a comprehensive literature review and propose the first two-dimensional evaluation taxonomy: one axis captures *objective dimensions*—including behavior, capability, reliability, and security—while the other captures *process dimensions*—encompassing interaction paradigms, benchmark datasets, evaluation metrics, and toolchains. Crucially, our taxonomy explicitly incorporates enterprise challenges, exposing critical shortcomings in existing work regarding holistic coverage, scalability, and real-world applicability. The framework provides researchers and practitioners with a structured, actionable reference for designing, evaluating, and deploying LLM agents in complex, mission-critical environments. It advances the field toward trustworthy, production-ready LLM agent systems. (138 words)
📝 Abstract
The rise of LLM-based agents has opened new frontiers in AI applications, yet evaluating these agents remains a complex and underdeveloped area. This survey provides an in-depth overview of the emerging field of LLM agent evaluation, introducing a two-dimensional taxonomy that organizes existing work along (1) evaluation objectives -- what to evaluate, such as agent behavior, capabilities, reliability, and safety -- and (2) evaluation process -- how to evaluate, including interaction modes, datasets and benchmarks, metric computation methods, and tooling. In addition to taxonomy, we highlight enterprise-specific challenges, such as role-based access to data, the need for reliability guarantees, dynamic and long-horizon interactions, and compliance, which are often overlooked in current research. We also identify future research directions, including holistic, more realistic, and scalable evaluation. This work aims to bring clarity to the fragmented landscape of agent evaluation and provide a framework for systematic assessment, enabling researchers and practitioners to evaluate LLM agents for real-world deployment.