🤖 AI Summary
This work addresses the fragmented and model-centric nature of existing evaluation methods for large language model (LLM) agents, which often overlook the influence of architectural components—such as planners, memory modules, and tool routers—on agent behavior, resulting in assessments that lack diagnostic precision and specificity. To bridge this gap, the paper proposes a lightweight, architecture-aware evaluation framework that systematically establishes the first explicit mapping between internal agent components, observable behaviors, and evaluation metrics. This approach shifts the paradigm from black-box assessment toward interpretable, component-level diagnosis. Through architecture-aware analysis, behavior-component modeling, and tailored metric design, the framework is validated on real-world LLM agents, demonstrating significant improvements in evaluation transparency, target specificity, and practical utility.
📝 Abstract
LLM-based agents are becoming central to software engineering tasks, yet evaluating them remains fragmented and largely model-centric. Existing studies overlook how architectural components, such as planners, memory, and tool routers, shape agent behavior, limiting diagnostic power. We propose a lightweight, architecture-informed approach that links agent components to their observable behaviors and to the metrics capable of evaluating them. Our method clarifies what to measure and why, and we illustrate its application through real world agents, enabling more targeted, transparent, and actionable evaluation of LLM-based agents.