benchmark auditing

Designs and executes audits of benchmarks, evaluation datasets, and automated evaluators to assess their validity, reliability, and failure modes. Builds and runs procedures that compare evaluator outputs to expert or reference judgments, perform sanity checks, quantify evaluator–human misalignment and run-to-run score variance, produce trace-level audit artifacts, and recommend benchmark refinements or formal dataset fixes.

benchmarkauditing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses a critical gap in existing tool-calling evaluation benchmarks: the lack of validation of the evaluators themselves, which risks conflating assessment artifacts with agents’ true capabilities. Through a systematic audit of four prominent benchmarks—BFCL v4, τ2-Bench, LiveMCPBench, and MCP-Atlas—the authors conduct expert review of 496 tasks, replicate experiments, and perform trajectory-level analysis, revealing an 18.5% disagreement rate between automated evaluators and human judgment. Notably, LiveMCPBench exhibits a score variance of up to 18.9 percentage points upon re-evaluation, sufficient to overturn leaderboard rankings. To address these issues, the work introduces the first unified taxonomy of tool-calling evaluation failures, advocates for distinct measurement of tool invocation, task completion, and result verification, and releases Tool-Veritas—a configurable benchmark—and Harness Lab, an open-source evaluation platform.

benchmark validityevaluator alignmentLLM benchmarks

Existing LLM-based evaluators exhibit low reliability for software engineering artifacts—such as code generation, translation, and summarization—due to high cost, subjectivity, and poor scalability of human evaluation, and insufficient sensitivity of automated methods to fine-grained quality differences. Method: We propose REFINE, a framework that integrates controllable fine-grained quality degradation synthesis with evaluator ranking alignment testing. It enables progressive tuning—from coarse-grained filtering to stress-testing subtle quality distinctions—supported by hierarchical dataset construction and a quantitative ranking consistency metric. Contribution/Results: Evaluated on industrial-scale COBOL code data, REFINE automatically identifies and validates high-quality evaluator configurations. Experiments show it improves alignment between LLM evaluators and human annotations from <0.7 to >0.9 (Kendall’s τ). The framework has been deployed to support model release decisions in production training teams.

Automating fine-grained quality assessment of code artifactsImproving reliability of LLM evaluators in coding tasksValidating LLM-based evaluators for software engineering artifacts

Eval Factsheets: A Structured Framework for Documenting AI Evaluations

Dec 03, 2025
FB
Florian Bordes
🏛️ FAIR | Meta

Current AI evaluation methodologies suffer from a lack of standardized, systematic documentation, severely undermining reproducibility, transparency, and trustworthy decision-making. Method: This paper introduces Eval Factsheets—a novel framework that pioneers the application of structured documentation to AI evaluation. It establishes a five-dimensional taxonomy—encompassing Context, Scope, Structure, Methodology, and Alignment—to uniformly characterize diverse evaluation paradigms, including traditional benchmarks and LLM-as-judge approaches. A taxonomy-guided questionnaire specifies mandatory and recommended fields covering the entire evaluation lifecycle. Contribution/Results: Empirical validation across multiple benchmark cases demonstrates that Eval Factsheets consistently represent heterogeneous evaluation practices, significantly enhancing cross-evaluation comparability, reproducibility, and transparency. The framework provides a foundational, extensible tool for standardizing AI evaluation documentation and practice.

Challenges in reproducibility and transparency due to benchmark proliferation.Lack of systematic documentation standards for AI evaluation methodologies.Need for structured framework to document diverse evaluation paradigms.

Current complex LLM agent benchmarks are often prone to misjudgment due to specification errors, implicit assumptions, or rigid evaluation scripts, mistakenly attributing benchmark flaws to agent failures. This work proposes the first automated auditing framework leveraging state-of-the-art large language models to cross-validate task-oriented, execution-driven benchmark components through structured LLM protocols, augmented by agent solutions and execution traces for diagnostic support. The approach establishes a novel AI-assisted paradigm for benchmark validation, overcoming the limitations of traditional manual review. Applied to ScienceAgentBench, it identified 12 author-confirmed issues—including critical errors—and reproduced 83.3% of expert-discovered problems on the BIXBench Verified-50 subset, with an auditing cost of under $15 for 50 bioinformatics tasks.

automated validationbenchmark auditingbenchmark flaws

Current evaluations of agent tool use often conflate workload specifications, action generation, and evidentiary criteria, lacking a unified and auditable framework. This work proposes an evaluation paradigm centered on “evidence admissibility gating,” which explicitly decouples workloads, drivers, and verification evidence through a shared evidence admissibility contract. The framework integrates diverse environments—including WebArena Verified, a subset of SWE-Gym, and MiniWoB++—and employs a standardized reporting pipeline comprising a universal workload adapter, declarative drivers, task manifests, event schemas, and replay/freeze strategies. It uniformly logs multidimensional metrics such as latency, invalid actions, and patching costs, enabling consistent differentiation of controller performance under identical workloads while ensuring relevance, reproducibility, and auditability in agent evaluations.

benchmarkingevaluation methodologyevidence admission

Latest Papers

What's happening recently
View more

Current agent benchmarks rely on manual auditing, which struggles to scale and often fails to identify validity flaws, thereby undermining the credibility of model capability evaluations. This work proposes the first automated AI scanner tailored for agent benchmarking, leveraging large language models and structured scoring rules to detect four categories of validity issues in agent transcripts. The system is calibrated through human annotation validation and cross-benchmark evaluation. Experiments across five prominent benchmarks uncover multiple quality issues that evade manual spot-checking, demonstrating that the proposed method effectively enables systematic auditing of benchmarks. This approach establishes a new paradigm for enhancing the reliability of agent evaluations.

agentic benchmarksautomated transcript analysisbenchmark auditing

Traditional AI evaluation methods are primarily designed for static model selection and often fail to diagnose root causes of performance degradation in production or guide targeted improvements. This work proposes EvalLoop, a novel methodology that embeds evaluation into a continuous optimization loop. By integrating dimensional metric grouping, failure mode categorization, single-variable controlled experiments, and a human-in-the-loop gating mechanism, EvalLoop enables precise mapping from failure attribution to actionable refinements and supports deployment-aware model selection. Evaluated on a sales intelligence briefing generation task, the approach increased overall accuracy of the best-performing model from 82.6% to 94.6%, improved performance on critical dimensions by over 16 percentage points, and reduced human review effort by 94%.

business AI systemsevaluationfailure diagnosis

Hot Scholars

AR

Anka Reuel

CS Ph.D. Candidate, Stanford University
AI GovernanceResponsible AIAI EthicsAI Safety
SK

Sanmi Koyejo

Assistant Professor, Stanford University
Machine LearningHealthcare AINeuroinformatics
OS

Olawale Salaudeen

Postdoctoral Associate, MIT
Trustworthy AIAI for SocietyDistribution ShiftCausality
LC

Lele Cao

Senior Principal AI/ML Researcher and Research Lead, Microsoft (ABK); ACM & IEEE Member
Machine LearningGraph LearningTime Series ModelingFinance
IS

Ion Stoica

Professor of Computer Science, UC Berkeley
Cloud ComputingNetworkingDistributed SystemsBig Data