cross-evaluation benchmarking

Designs and implements evaluation benchmarks and protocols that decouple judge systems from target models to enable cross-evaluation: aggregating judgments from multiple automated or human judges, comparing those judgments to subject-matter-expert ground truth, and producing standardized metrics and rankings of model reliability across tasks. Builds cross-evaluation frameworks and benchmark suites and analyzes aggregated evaluation outputs to identify failure modes, measure inter-model agreement, and support reproducible ranking and calibration of models.

cross-evaluationbenchmarking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.65
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses a critical gap in existing tool-calling evaluation benchmarks: the lack of validation of the evaluators themselves, which risks conflating assessment artifacts with agents’ true capabilities. Through a systematic audit of four prominent benchmarks—BFCL v4, τ2-Bench, LiveMCPBench, and MCP-Atlas—the authors conduct expert review of 496 tasks, replicate experiments, and perform trajectory-level analysis, revealing an 18.5% disagreement rate between automated evaluators and human judgment. Notably, LiveMCPBench exhibits a score variance of up to 18.9 percentage points upon re-evaluation, sufficient to overturn leaderboard rankings. To address these issues, the work introduces the first unified taxonomy of tool-calling evaluation failures, advocates for distinct measurement of tool invocation, task completion, and result verification, and releases Tool-Veritas—a configurable benchmark—and Harness Lab, an open-source evaluation platform.

benchmark validityevaluator alignmentLLM benchmarks

JudgeBench: A Benchmark for Evaluating LLM-based Judges

Oct 16, 2024
ST
Sijun Tan
🏛️ UC Berkeley | Washington University in St. Louis

Existing LLM judge evaluation benchmarks inadequately assess judges’ ability to discern factual accuracy and logical correctness in knowledge, reasoning, mathematics, and programming tasks. Method: We introduce the first objective-correctness–oriented LLM judge benchmark, featuring an automated pipeline for response pairing and preference labeling derived from diverse challenging sources—including MMLU, GSM8K, and HumanEval—where ground-truth factual and logical correctness serves as the sole, verifiable evaluation criterion, eliminating reliance on human preferences. Contribution/Results: Our framework enables rigorous evaluation of strong judge models (e.g., GPT-4o) across mainstream paradigms: prompt engineering, fine-tuning, multi-agent systems, and reward modeling. Experiments reveal that state-of-the-art judge models achieve only ~55% accuracy—substantially below human performance—demonstrating the benchmark’s high difficulty and validity in exposing critical limitations in current judge capabilities.

Assessing judges on challenging tasks beyond human preferencesCreating a benchmark for advanced LLM judge evaluationEvaluating reliability of LLM-based judges objectively

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities

Dec 09, 2024
AG
Adhiraj Ghosh
🏛️ University of Tübingen | Open-Ψ (Open-Sci) Collective | University of Cambridge

Traditional static test sets inadequately evaluate foundation models’ diverse capabilities in open-ended scenarios. To address this, we propose ONEBench—a dynamic, extensible benchmarking paradigm that enables on-demand generation of customized evaluation suites targeting open capabilities, framing model assessment as a collective selection and aggregation process over sample-level tests. Our key contributions include: (1) the first unified, open-ended, and evolvable evaluation framework operating at the sample level; (2) a sparse measurement aggregation algorithm, a progressive sample pool construction mechanism, and a cross-modal unified interface (ONEBench-LLM/LMM); and (3) a robustness-aware scoring model with theoretical guarantees on identifiability and fast convergence. Experiments show that ONEBench achieves ranking stability >0.98 under 95% measurement sparsity, reduces evaluation cost by 20×, and attains >0.98 correlation with mean-score rankings on homogeneous data—enabling unified, efficient, and reliable assessment of both language and multimodal models.

Aggregating diverse metrics into reliable model scoresEvaluating open-ended capabilities of foundation modelsReducing evaluation cost while maintaining accuracy

This work addresses the proliferation of large language model (LLM) evaluation benchmarks, which has outpaced systematic assessment of their intrinsic quality. To this end, we propose Benchmark², a novel framework that establishes the first quantitative methodology for evaluating the reliability and validity of LLM benchmarks through three complementary metrics: cross-benchmark ranking consistency, discriminability score, and capability alignment bias. Empirical evaluation across 15 benchmarks and 11 LLMs demonstrates that Benchmark² not only reveals substantial quality disparities among existing benchmarks but also enables the construction of streamlined test sets that maintain high evaluative performance while significantly reducing assessment scale.

benchmark qualitybenchmark reliabilityLLM benchmarks

Latest Papers

What's happening recently
View more

Current evaluations of AI models lack standardized protocols, with institutions selectively employing benchmarks in ways that hinder cross-study comparability and raise concerns about scientific validity. This work introduces Benchmarking-Cultures-25, a dataset encompassing 231 benchmarks from 139 model releases, and combines qualitative content analysis with a unified categorization framework to systematically expose the fragmentation in benchmark selection: 63.2% of benchmarks are used by only a single institution, and 38.5% appear just once. Moreover, many benchmarks marketed as “general-purpose” disproportionately emphasize STEM—particularly mathematics—while often neglecting construct validity. The study further proposes a taxonomy aligning ostensibly disparate terminologies to their underlying measurement signals and develops an interactive tool revealing that benchmarks frequently serve marketing narratives rather than rigorous scientific assessment.

AI evaluationbenchmarkingconstruct validity

Alignment evaluation in machine learning has largely become evaluation of models. Influential benchmarks score model outputs under fixed inputs, such as truthfulness, instruction following, or pairwise preference, and these scores are often used to support claims about deployed alignment. This paper argues that deployment-relevant alignment cannot be inferred from model-level evaluation alone. Alignment claims should instead be indexed to the level at which evidence is collected: model-level, response-level, interaction-level, or deployment-level. Two studies support this position. First, a structured audit of eleven alignment benchmarks, extended to a sixteen-benchmark corpus, dual-coded against an eight-dimension rubric with Cohen's kappa = 0.87, finds that user-facing verification support is absent across every benchmark examined, while process steerability is nearly absent. The few interactional benchmarks identified, including tau-bench, CURATe, Rifts, and Common Ground, remain fragmented in coverage, and benchmark construction rather than data source determines what is measured. Second, a blinded cross-model stress test using 180 transcripts across three frontier models and four scaffolds finds that the same verification scaffold raises one model's verification support to ceiling while leaving another categorically unchanged. This shows that scaffold efficacy is model-dependent and that the gap identified by the audit cannot be closed at the model level alone. We propose a system-level evaluation agenda: alignment profiles instead of single scores, fixed-scaffolding protocols for comparable interactional evaluation, and reporting templates that make the inferential distance between evaluation evidence and deployment claims explicit.

alignment evaluationbenchmark limitationsdeployment-relevant alignment

In highly specialized domains lacking ground-truth answers, it remains unclear how to effectively evaluate the quality of generated content and whether scoring rubrics or pairwise preferences constitute more suitable supervision signals. This work introduces JudgmentBench, a benchmark comprising 30 real-world legal tasks, where the same cohort of experienced lawyers provides both rubric-based scores and pairwise preference annotations for three quality tiers of outputs generated by large language models. Empirical analysis demonstrates that pairwise preference judgments substantially outperform rubric-based scoring in both validity—evidenced by a Spearman correlation coefficient of 0.908 versus 0.150—and efficiency, requiring less than half the annotation time. These findings hold consistently across both human and automated evaluators. The study delivers the first dual-modality expert-annotated dataset and methodological foundation for evaluation in high-expertise domains.

benchmarkingcomparative judgmentexpert judgment

Existing foundational model evaluation benchmarks often suffer from insufficient fine-grained coverage and a lack of metadata, limiting their ability to comprehensively assess model capabilities. This work proposes an automated benchmark generation framework that uniquely integrates a multi-agent system with a solution-graph-driven question-generation mechanism to synthesize high-quality, contamination-resistant evaluation items from reference materials such as textbooks, while automatically annotating fine-grained metadata. The approach substantially reduces ground-truth error rates and achieves near-uniform coverage across capability dimensions. Leveraging this framework, we construct three new benchmarks in machine learning, corporate finance, and personal finance. Expert evaluations demonstrate that these benchmarks exhibit significantly lower error rates than MMLU and GSM8K and uncover performance differences among models that existing benchmarks fail to detect.

benchmark generationcomprehensive coveragefine-grained evaluation

Hot Scholars

YW

Yichao Wu

SenseTime Group Limited
AGILLMComputer VisionFace recognition
YY

Yifan Yang

Senior Research SDE, Microsoft Research Asia
Multi-modalityComputer VisionMachine LearningArtificial Intelligence
HZ

Hongyu Zhang

Chongqing University
Software EngineeringMining Software RepositoriesData-driven Software EngineeringSoftware Analytics
ZC

Ziyan Chen

Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences
Generative AILow Level Vision
NB

Nolwenn Bernard

TH Köln
User SimulationConversational Information AccessNatural Language Processing