Score
Design and implement evaluation protocols, benchmarks, and pipelines for interactive, multi-turn agents — including session-based and replayed-session tests, task-driven and perception-driven evaluations, and multi-agent simulation setups — and build tooling to run interactive diagnosis and benchmarking. Analyze outcomes and failure modes by measuring metrics such as chain completion and final-task correctness, counts of corrective-feedback turns and interventions, session management and engagement, and differences between single-turn and multi-turn protocols.
Current LLM agent evaluation lacks a systematic framework, particularly neglecting enterprise-specific requirements such as role-based access control, regulatory compliance, and long-horizon interactive behavior. To address this gap, we conduct a comprehensive literature review and propose the first two-dimensional evaluation taxonomy: one axis captures *objective dimensions*—including behavior, capability, reliability, and security—while the other captures *process dimensions*—encompassing interaction paradigms, benchmark datasets, evaluation metrics, and toolchains. Crucially, our taxonomy explicitly incorporates enterprise challenges, exposing critical shortcomings in existing work regarding holistic coverage, scalability, and real-world applicability. The framework provides researchers and practitioners with a structured, actionable reference for designing, evaluating, and deploying LLM agents in complex, mission-critical environments. It advances the field toward trustworthy, production-ready LLM agent systems. (138 words)
This study investigates whether increasing the number of agents genuinely enhances the performance of large language model (LLM) workflows under a unified evaluation protocol. To this end, we introduce BenchAgent, a novel evaluation framework that establishes a protocol-aligned standardization paradigm, enabling fair comparisons among single-agent, fixed multi-agent, and evolving multi-agent workflows under identical conditions—including benchmark loaders, tool access, answer contracts, and trajectory logging. Experiments across ten reasoning, programming, and tool-use benchmarks reveal that among six multi-agent systems, only EvoAgent approaches the performance of single-agent systems, while the others lag by 2.56–11.29 percentage points. Notably, runtime-generated workflows based on GPT-4.1 and Claude-Code architectures achieve 66.72% accuracy on GAIA, significantly outperforming fixed multi-agent baselines.
Current AI evaluation predominantly relies on static responses, which inadequately capture the systematic capabilities of large language models in dynamic scenarios involving tool use, environmental interaction, and multi-agent collaboration. This work proposes “interactive evaluation” as a distinct paradigm, introducing a two-axis taxonomy to clarify its design principles and reporting standards. By leveraging trajectory modeling and a multi-dimensional scoring mechanism, the framework extracts evidence from interaction processes to holistically assess model performance across dimensions such as procedural fidelity, recoverability, coordination, robustness, and system-level efficacy. The study delineates core challenges in interactive evaluation and redefines the logical mapping from empirical evidence to performance judgments, thereby establishing a theoretical foundation for a unified, comparable, and interpretable next-generation AI evaluation framework.
Existing evaluation paradigms for interactive agents overemphasize task success rate while neglecting holistic user experience across the entire interaction process. Method: This paper proposes the PIPA protocol, the first framework to model interactive task-planning agent behavior as a Partially Observable Markov Decision Process (POMDP). It decomposes the agent’s behavioral chain into atomic components—context understanding, tool invocation, and response generation—and introduces multi-granularity, interpretable evaluation metrics. Contribution/Results: By correlating intermediate behavioral steps with user satisfaction, PIPA uncovers uneven capability distributions across stages and empirically validates the substantial impact of intermediate behaviors on end-to-end user experience. It delivers actionable, fine-grained diagnostic insights to guide agent optimization and identifies concrete directions for advancing multi-agent coordination and user simulator design.
Existing benchmarks for interactive agents struggle to simultaneously ensure scalability and effectively evaluate performance under realistic workflow conditions involving state conflicts—such as partial, stale, or contradictory prior states. To address this gap, this work proposes ClawForge, the first executable command-line workflow framework that systematically supports evaluation under pre-existing state conflicts. ClawForge enables reproducible task construction through scenario templates, state initialization, reference trajectories, and validators, while abandoning strict trajectory matching in favor of stepwise assessment based on normalized final states and observable side effects. The accompanying ClawForge-Bench comprises 17 scenarios; evaluations reveal that even the best-performing model achieves only a 45.3% strict accuracy, with all models exhibiting errorful state replacement rates below 17%. Critically, the tendency to proactively inspect existing states emerges as a key determinant of performance disparities.
Existing evaluation of generative AI is hindered by the scarcity of high-quality benchmarks, whose manual construction is costly and time-consuming. Method: We propose the first automated benchmark construction framework powered by collaborative large language model (LLM) agents, decomposing benchmark creation into four sequential stages—planning, generation, verification, and evaluation—integrating task decomposition, agent coordination, human-in-the-loop feedback, and explicit constraint-satisfaction assessment. Contribution/Results: The framework significantly enhances data diversity and metric reliability. Leveraging it, we construct the first high-quality benchmark specifically targeting planning and constraint-satisfaction capabilities in text generation. We systematically evaluate seven state-of-the-art models, uncovering shared failure modes and fine-grained capability disparities. Our work establishes a scalable, reproducible paradigm for evaluating generative AI capabilities, advancing both benchmark methodology and empirical analysis.
Current agent evaluation lacks open, general, and reproducible interfaces, leading to a disconnect between benchmarking and real-world deployment. This work proposes the Agentic Agent Assessment (AAA) framework, which introduces evaluator agents to conduct assessments solely through a unified Agent-to-Agent (A2A) task protocol and Model Calling Protocol (MCP), enabling interoperable, reproducible, and multi-agent collaborative evaluation of heterogeneous agents. The framework supports five operational modes that balance openness, privacy, and practicality. Its advantages in evaluation breadth, utility, and fidelity are demonstrated through a five-month open competition involving 298 evaluator agents and 467 participant agents, as well as a case study on programming agents.
This work addresses the limitation of existing benchmarks, which focus solely on accuracy in multi-agent orchestration tasks while neglecting fine-grained diagnosis of failure origins and recovery capabilities. The authors propose a reproducible fault-injection framework to systematically evaluate failure modes, task decomposition quality, and recovery mechanisms within templated enterprise workflows. They introduce two novel metrics: “cascade radius” and failure-mode-specific recovery rates, and employ controlled probes to analyze recovery behavior across different fault types. Experimental results demonstrate that intent-based reasoning routing achieves 100% recovery under adversarial conditions, significantly outperforming keyword-based routing; tool-related failures are fully recoverable, whereas semantic failures prove largely irrecoverable; and cascade radius increases with workflow depth.
Current agent benchmarks often yield misleading evaluation scores due to invalid protocols, primarily stemming from reward hacking or assessment vulnerabilities. This work presents the first systematic formalization of “protocol validity” and introduces Mislead Gap—a quantitative metric—and HackDetect, a posterior auditing framework. By integrating trajectory auditing, exposure point identification, and intent-exploitation score comparison, the framework uniformly detects and quantifies the impact of reward hacking. Empirical analysis across 15 benchmarks and 2,385 agent trajectories reveals that 66.7%–67.0% of evaluations exhibit exposure to or active engagement in reward hacking, inflating scores by 0.45–1.00. These findings demonstrate that prevailing benchmarks generally fail to validate agents’ true capabilities.
Current evaluation methods for personal agents assess isolated dimensions—such as memory, tool use, or safety—independently, failing to capture the evolving state dynamics and cross-component failure propagation inherent in real-world user interactions over time. This work proposes a user-conditioned state evaluation protocol that formally defines four necessary criteria for meaningful personal agent assessment: explicit temporal intervention, state persistence, cross-dimensional effect induction, and user-conditioned state change. Through focused benchmark auditing, formal modeling, and minimalistic design, the study identifies critical gaps in existing benchmarks and constructs a minimal evaluation framework that satisfies all four criteria, accompanied by a corresponding metric suite. This provides a clear and principled technical pathway for future evaluations of personal agents.
Current computer-using agent (CUA) benchmarks rely on fragile scripted evaluators that frequently produce erroneous failure judgments, obscuring true performance bottlenecks. This work proposes the first reliability-focused evaluation framework encompassing the entire pipeline—from task construction and trajectory observation to scoring and reporting—and introduces a three-tier failure diagnosis taxonomy. Through manual auditing and attribution analysis of 150 publicly reported failure trajectories, we find that 15.3% of failure labels are incorrect, with 10.7% stemming from evaluator misjudgment and 4.7% arising from task design flaws. Building on these insights, we derive phased design principles for long-horizon CUA evaluation, substantially improving assessment accuracy and interpretability.