llm agent benchmarking

Designs, builds, and runs evaluation suites and benchmarking platforms that execute LLM agents and agent pipelines in controlled, reproducible scenarios—creating task benchmarks, executable scenario trajectories, preserved frozen-prompt inputs and provider-form outputs, and curated collections of multimodal tasks organized by difficulty. Defines metrics, ground-truth labels, and protocols to measure functional, security, retrieval-screening-synthesis and economic outcomes, and implements isolated evaluation environments, feedback-tiered runs, simulators and comparison frameworks to produce reproducible agent evaluations and baselined comparisons.

llmagentbenchmarking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Survey on Evaluation of LLM-based Agents

Mar 20, 2025
AY
Asaf Yehudai
🏛️ IBM Research | Yale University

This paper addresses the lack of a systematic framework for evaluating LLM-based agents. We propose the first four-dimensional taxonomy—encompassing foundational capabilities, domain-specific applications, general-purpose benchmarks, and evaluation frameworks—derived from a systematic literature review and multidimensional modeling of over one hundred empirical evaluation practices. Our analysis reveals an emerging trend toward realism and dynamism in agent evaluation, while identifying critical gaps in cost-efficiency, safety, robustness, and fine-grained scalable assessment. The main contributions are: (1) the first comprehensive, multi-dimensional taxonomy for LLM agent evaluation; (2) a holistic evaluation landscape map that clarifies current limitations; and (3) six concrete, actionable research directions to advance standardized, trustworthy agent evaluation. This work provides theoretical foundations for rigorous, reproducible, and application-aware assessment methodologies in the evolving field of LLM agents.

Analysis of benchmarks across planning, tool use, and memory.Comprehensive survey of LLM-based agent evaluation methodologies.Identification of gaps in cost-efficiency, safety, and robustness evaluation.

Must-Read Papers

Most classic and influential ideas
View more

Continuous Benchmark Generation for Evaluating Enterprise-scale LLM Agents

Nov 13, 2025
DS
Divyanshu Saxena
🏛️ The University of Texas at Austin | Microsoft

Enterprise-scale LLM agents face persistent evaluation challenges due to dynamic service evolution and scarcity of realistic, annotated test cases. Method: This paper proposes a dynamic benchmark generation approach grounded in semi-structured documentation, leveraging intent extraction and LLM-driven test case synthesis to automatically construct maintainable, business-aligned evaluation benchmarks that evolve with operational requirements—without reliance on dense human annotation and robust even under sparse ground-truth conditions. Contribution/Results: Compared to static benchmarks, our method significantly reduces maintenance overhead while improving coverage and responsiveness. Empirical validation in large-scale enterprise service migration scenarios demonstrates a 3.2× improvement in benchmark construction efficiency, enabling rapid agent iteration and closed-loop feedback for continuous optimization.

Creating maintainable evaluation frameworks using minimal semi-structured documentsEvaluating evolving enterprise-scale AI agents with sparse ground-truth dataGenerating continuous benchmarks to match changing service requirements

BENCHAGENTS: Automated Benchmark Creation with Agent Interaction

Oct 29, 2024
NB
Natasha Butt
🏛️ University of Amsterdam | Microsoft Research | UIUC

Existing evaluation of generative AI is hindered by the scarcity of high-quality benchmarks, whose manual construction is costly and time-consuming. Method: We propose the first automated benchmark construction framework powered by collaborative large language model (LLM) agents, decomposing benchmark creation into four sequential stages—planning, generation, verification, and evaluation—integrating task decomposition, agent coordination, human-in-the-loop feedback, and explicit constraint-satisfaction assessment. Contribution/Results: The framework significantly enhances data diversity and metric reliability. Leveraging it, we construct the first high-quality benchmark specifically targeting planning and constraint-satisfaction capabilities in text generation. We systematically evaluate seven state-of-the-art models, uncovering shared failure modes and fine-grained capability disparities. Our work establishes a scalable, reproducible paradigm for evaluating generative AI capabilities, advancing both benchmark methodology and empirical analysis.

Automating high-quality benchmark creation for evolving AI modelsGenerating structured benchmarks for complex reasoning and multimodal evaluationOvercoming slow manual benchmark creation via multi-agent framework

This work addresses the limitations of existing large model evaluation benchmarks—namely high construction costs, poor reusability, and rapid saturation—which hinder their ability to continuously differentiate the performance of state-of-the-art models. To overcome these challenges, the authors propose Benchmark Agent, the first end-to-end fully automated agent system capable of generating high-quality benchmarks with minimal human intervention. By integrating agent-based architecture, LLM-as-a-judge evaluation, human feedback, and consistency verification, the system autonomously handles query parsing, subtask design, data annotation, and quality control. It supports diverse evaluation scenarios and has successfully constructed 15 benchmarks spanning text understanding, multimodal comprehension, and domain-specific reasoning, effectively exposing models’ weaknesses in complex reasoning and demonstrating strong efficiency and scalability.

benchmarksevaluationlarge language models

Current complex LLM agent benchmarks are often prone to misjudgment due to specification errors, implicit assumptions, or rigid evaluation scripts, mistakenly attributing benchmark flaws to agent failures. This work proposes the first automated auditing framework leveraging state-of-the-art large language models to cross-validate task-oriented, execution-driven benchmark components through structured LLM protocols, augmented by agent solutions and execution traces for diagnostic support. The approach establishes a novel AI-assisted paradigm for benchmark validation, overcoming the limitations of traditional manual review. Applied to ScienceAgentBench, it identified 12 author-confirmed issues—including critical errors—and reproduced 83.3% of expert-discovered problems on the BIXBench Verified-50 subset, with an auditing cost of under $15 for 50 bioinformatics tasks.

automated validationbenchmark auditingbenchmark flaws

This work addresses the limitations of existing agent benchmarks, which often focus on isolated capabilities and struggle to evaluate long-horizon, high-complexity real-world tasks due to reliance on manual feedback that hinders scalability. The authors propose the first comprehensive, automated benchmark tailored to everyday AI usage scenarios, encompassing 32 real-world settings and 138 tasks—each requiring an average of 90 tool invocations and processing over one million tokens. The framework employs user-simulation agents for iterative feedback, Docker-based sandboxing for visual and functional rule validation, and a standardized task interface enabling unified closed-loop evaluation of both open- and closed-source models. Experimental results demonstrate a significant performance gap favoring closed-source models (48.4% vs. 32.1%) and highlight the critical role of co-optimizing models with agent frameworks to enhance overall effectiveness.

automated evaluationautonomous agentsbenchmarking

Latest Papers

What's happening recently
View more

Current evaluations of large language model (LLM) agents are often confounded by implementation-specific details and environmental variability, hindering fair assessment of intrinsic model capabilities. This work proposes a unified evaluation framework that standardizes diverse benchmarks into a consistent instruction–tool–environment format, executed within a controlled sandbox under a fixed ReAct architecture and supported by offline snapshots to decouple environmental influences. For the first time, cross-benchmark standardized evaluation is achieved, introducing unified metrics for resource consumption and a failure attribution taxonomy distinguishing decision-making from execution errors. The framework integrates seven major benchmarks spanning 24 domains, encompassing 400,000 rollouts and 5 billion tokens, revealing significant impacts of architectural and environmental factors on performance and effectively isolating true model capabilities from external interference.

benchmark evaluationcross-benchmark comparisonevaluation framework

This study investigates whether increasing the number of agents genuinely enhances the performance of large language model (LLM) workflows under a unified evaluation protocol. To this end, we introduce BenchAgent, a novel evaluation framework that establishes a protocol-aligned standardization paradigm, enabling fair comparisons among single-agent, fixed multi-agent, and evolving multi-agent workflows under identical conditions—including benchmark loaders, tool access, answer contracts, and trajectory logging. Experiments across ten reasoning, programming, and tool-use benchmarks reveal that among six multi-agent systems, only EvoAgent approaches the performance of single-agent systems, while the others lag by 2.56–11.29 percentage points. Notably, runtime-generated workflows based on GPT-4.1 and Claude-Code architectures achieve 66.72% accuracy on GAIA, significantly outperforming fixed multi-agent baselines.

agent collaborationbenchmarkingLLM agents

This work addresses the frequent failure of research code deployment due to complex environment configurations, heterogeneous toolchains, system dependencies (e.g., GPU/CUDA), and legacy compatibility issues—challenges inadequately captured by existing benchmarks. To bridge this gap, we introduce DeployBench, the first systematic, multidimensional deployment benchmark encompassing 51 research artifacts across AI/ML, computer systems, and scientific computing. DeployBench evaluates the autonomous deployment capabilities of LLM agents through hidden validation pipelines that reproduce experiments and verify outputs. Built upon the OpenHands framework, our evaluation integrates four state-of-the-art LLMs and executes end-to-end deployment and validation in full system environments. Results reveal that even the best-performing agent achieves a success rate of only 7.8%–51.0%, with 63% of failures attributed to premature self-termination or misaligned validation objectives, exposing critical deficiencies in agents’ judgment of task completion.

benchmarkenvironment setupLLM agents

This study addresses a critical gap in existing tool-calling evaluation benchmarks: the lack of validation of the evaluators themselves, which risks conflating assessment artifacts with agents’ true capabilities. Through a systematic audit of four prominent benchmarks—BFCL v4, τ2-Bench, LiveMCPBench, and MCP-Atlas—the authors conduct expert review of 496 tasks, replicate experiments, and perform trajectory-level analysis, revealing an 18.5% disagreement rate between automated evaluators and human judgment. Notably, LiveMCPBench exhibits a score variance of up to 18.9 percentage points upon re-evaluation, sufficient to overturn leaderboard rankings. To address these issues, the work introduces the first unified taxonomy of tool-calling evaluation failures, advocates for distinct measurement of tool invocation, task completion, and result verification, and releases Tool-Veritas—a configurable benchmark—and Harness Lab, an open-source evaluation platform.

benchmark validityevaluator alignmentLLM benchmarks

Hot Scholars

YZ

Yue Zhao

Assistant Professor of Computer Science, University of Southern California
Anomaly DetectionOut-of-Distribution DetectionTrustworthy AIAI for Science
YM

Yuchen Ma

LMU Munich
Deep LearningCausal InferenceDiffusion ModelsFoundation Models
SM

Sicheng Mo

University of California, Los Angeles
Computer Vision
TY

Toshihiko Yamasaki

Department of Information and Communication Engineering, The University of Tokyo
Image ProcessingMultimediaComputer VisionPattern Recognition