Score
Designs and executes audits of benchmarks, evaluation datasets, and automated evaluators to assess their validity, reliability, and failure modes. Builds and runs procedures that compare evaluator outputs to expert or reference judgments, perform sanity checks, quantify evaluator–human misalignment and run-to-run score variance, produce trace-level audit artifacts, and recommend benchmark refinements or formal dataset fixes.
This study addresses a critical “evaluation–safety gap” (EvalSafetyGap) in large language models (LLMs), wherein apparent performance gains do not necessarily reflect genuine safety capabilities. Through a systematic literature review, gray literature analysis, and a multidimensional audit of ten models across eight evidence streams, the work proposes the EvalSafetyGap hypothesis and introduces two novel constructs—“instability decomposition” and the “alignment trilemma”—to establish a unified terminology and evidence map supporting dynamic evaluation and auditable alignment. Empirical findings reveal no significant correlation between model capability and adversarial robustness (r = 0.232, p = 0.520). Moreover, safety differences between open- and closed-source models stem primarily from governance transparency rather than behavioral robustness, with results highly sensitive to model categorization and evaluation protocols.
This study addresses a critical gap in existing tool-calling evaluation benchmarks: the lack of validation of the evaluators themselves, which risks conflating assessment artifacts with agents’ true capabilities. Through a systematic audit of four prominent benchmarks—BFCL v4, τ2-Bench, LiveMCPBench, and MCP-Atlas—the authors conduct expert review of 496 tasks, replicate experiments, and perform trajectory-level analysis, revealing an 18.5% disagreement rate between automated evaluators and human judgment. Notably, LiveMCPBench exhibits a score variance of up to 18.9 percentage points upon re-evaluation, sufficient to overturn leaderboard rankings. To address these issues, the work introduces the first unified taxonomy of tool-calling evaluation failures, advocates for distinct measurement of tool invocation, task completion, and result verification, and releases Tool-Veritas—a configurable benchmark—and Harness Lab, an open-source evaluation platform.
Existing LLM-based evaluators exhibit low reliability for software engineering artifacts—such as code generation, translation, and summarization—due to high cost, subjectivity, and poor scalability of human evaluation, and insufficient sensitivity of automated methods to fine-grained quality differences. Method: We propose REFINE, a framework that integrates controllable fine-grained quality degradation synthesis with evaluator ranking alignment testing. It enables progressive tuning—from coarse-grained filtering to stress-testing subtle quality distinctions—supported by hierarchical dataset construction and a quantitative ranking consistency metric. Contribution/Results: Evaluated on industrial-scale COBOL code data, REFINE automatically identifies and validates high-quality evaluator configurations. Experiments show it improves alignment between LLM evaluators and human annotations from <0.7 to >0.9 (Kendall’s τ). The framework has been deployed to support model release decisions in production training teams.
Current AI evaluation methodologies suffer from a lack of standardized, systematic documentation, severely undermining reproducibility, transparency, and trustworthy decision-making. Method: This paper introduces Eval Factsheets—a novel framework that pioneers the application of structured documentation to AI evaluation. It establishes a five-dimensional taxonomy—encompassing Context, Scope, Structure, Methodology, and Alignment—to uniformly characterize diverse evaluation paradigms, including traditional benchmarks and LLM-as-judge approaches. A taxonomy-guided questionnaire specifies mandatory and recommended fields covering the entire evaluation lifecycle. Contribution/Results: Empirical validation across multiple benchmark cases demonstrates that Eval Factsheets consistently represent heterogeneous evaluation practices, significantly enhancing cross-evaluation comparability, reproducibility, and transparency. The framework provides a foundational, extensible tool for standardizing AI evaluation documentation and practice.
Current complex LLM agent benchmarks are often prone to misjudgment due to specification errors, implicit assumptions, or rigid evaluation scripts, mistakenly attributing benchmark flaws to agent failures. This work proposes the first automated auditing framework leveraging state-of-the-art large language models to cross-validate task-oriented, execution-driven benchmark components through structured LLM protocols, augmented by agent solutions and execution traces for diagnostic support. The approach establishes a novel AI-assisted paradigm for benchmark validation, overcoming the limitations of traditional manual review. Applied to ScienceAgentBench, it identified 12 author-confirmed issues—including critical errors—and reproduced 83.3% of expert-discovered problems on the BIXBench Verified-50 subset, with an auditing cost of under $15 for 50 bioinformatics tasks.
Current evaluations of agent tool use often conflate workload specifications, action generation, and evidentiary criteria, lacking a unified and auditable framework. This work proposes an evaluation paradigm centered on “evidence admissibility gating,” which explicitly decouples workloads, drivers, and verification evidence through a shared evidence admissibility contract. The framework integrates diverse environments—including WebArena Verified, a subset of SWE-Gym, and MiniWoB++—and employs a standardized reporting pipeline comprising a universal workload adapter, declarative drivers, task manifests, event schemas, and replay/freeze strategies. It uniformly logs multidimensional metrics such as latency, invalid actions, and patching costs, enabling consistent differentiation of controller performance under identical workloads while ensuring relevance, reproducibility, and auditability in agent evaluations.
Current agent benchmarks rely on manual auditing, which struggles to scale and often fails to identify validity flaws, thereby undermining the credibility of model capability evaluations. This work proposes the first automated AI scanner tailored for agent benchmarking, leveraging large language models and structured scoring rules to detect four categories of validity issues in agent transcripts. The system is calibrated through human annotation validation and cross-benchmark evaluation. Experiments across five prominent benchmarks uncover multiple quality issues that evade manual spot-checking, demonstrating that the proposed method effectively enables systematic auditing of benchmarks. This approach establishes a new paradigm for enhancing the reliability of agent evaluations.
Traditional AI evaluation methods are primarily designed for static model selection and often fail to diagnose root causes of performance degradation in production or guide targeted improvements. This work proposes EvalLoop, a novel methodology that embeds evaluation into a continuous optimization loop. By integrating dimensional metric grouping, failure mode categorization, single-variable controlled experiments, and a human-in-the-loop gating mechanism, EvalLoop enables precise mapping from failure attribution to actionable refinements and supports deployment-aware model selection. Evaluated on a sales intelligence briefing generation task, the approach increased overall accuracy of the best-performing model from 82.6% to 94.6%, improved performance on critical dimensions by over 16 percentage points, and reduced human review effort by 94%.