benchmark creation

Designs and builds end-to-end benchmarking artifacts and frameworks: curated datasets and real-world corpora (including community-annotated or consortium-developed resources), annotation schemas and multi-level taxonomies, label sets from binary to fine-grained, balanced category distributions, and multi-turn or open-response model–response collections with human validation. Specifies train/validation/test splits and forbidden-feature/label rules, implements expert-in-the-loop annotation workflows and annotator training, and defines evaluation protocols, metrics, baselines, and measurement procedures to evaluate diverse model families across specified performance dimensions.

benchmarkcreation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.87
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$213K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

BENCHAGENTS: Automated Benchmark Creation with Agent Interaction

Oct 29, 2024
NB
Natasha Butt
🏛️ University of Amsterdam | Microsoft Research | UIUC

Existing evaluation of generative AI is hindered by the scarcity of high-quality benchmarks, whose manual construction is costly and time-consuming. Method: We propose the first automated benchmark construction framework powered by collaborative large language model (LLM) agents, decomposing benchmark creation into four sequential stages—planning, generation, verification, and evaluation—integrating task decomposition, agent coordination, human-in-the-loop feedback, and explicit constraint-satisfaction assessment. Contribution/Results: The framework significantly enhances data diversity and metric reliability. Leveraging it, we construct the first high-quality benchmark specifically targeting planning and constraint-satisfaction capabilities in text generation. We systematically evaluate seven state-of-the-art models, uncovering shared failure modes and fine-grained capability disparities. Our work establishes a scalable, reproducible paradigm for evaluating generative AI capabilities, advancing both benchmark methodology and empirical analysis.

Automating high-quality benchmark creation for evolving AI modelsGenerating structured benchmarks for complex reasoning and multimodal evaluationOvercoming slow manual benchmark creation via multi-agent framework

Existing multimodal evaluation benchmarks inadequately reflect real-world, heterogeneous daily usage scenarios and lack systematic assessment across diverse tasks and output formats. Method: We introduce the first fine-grained, real-scenario-oriented multimodal benchmark—comprising 505 practical scenarios and 8,000+ samples—supporting 16 input/output modalities and 40+ output formats (e.g., numbers, code, JSON, free-form text). We propose a four-dimensional capability reporting framework—“Application–Input–Output–Skill”—replacing monolithic multiple-choice evaluation with task-driven, format-aware, interpretable assessment. The benchmark integrates expert crowdsourced scenario sampling, 40+ customized automated metrics, multi-format parsers, and interactive visualization tools. Contribution/Results: Comprehensive evaluation of state-of-the-art vision-language models reveals, for the first time, their fine-grained capability boundaries and long-tail deficiencies across modality combinations and task types.

Evaluates models with 40+ metrics across varied output formatsOptimizes diverse high-quality data for cost-effective evaluationScales multimodal evaluation to 500+ real-world tasks

DCA-Bench: A Benchmark for Dataset Curation Agents

Jun 11, 2024
BH
Benhao Huang
🏛️ Shanghai Jiao Tong University | University of Illinois Urbana-Champaign | University of Michigan

Dataset quality defects—such as missing documentation, incorrect labels, and ethical risks—are pervasive in open platforms yet resistant to detection by rule-based scripts, necessitating intelligent, automated identification methods. Method: We introduce the first LLM-agent benchmark for discovering real-world dataset quality issues, covering 221 empirically validated cases across eight platforms. It uniquely evaluates agents’ ability to autonomously detect latent defects without prior prompting. We propose an automated evaluation framework powered by GPT-4o, achieving high agreement with human experts (Cohen’s κ = 0.89), and ensure benchmark reliability via multi-source real-data sampling and expert annotation. Contribution/Results: Experiments reveal that even the state-of-the-art Curator agent detects only ~30% of defects, underscoring task difficulty. All benchmark data, code, and evaluation tools are publicly released to advance intelligent data governance.

Assessing dataset quality issues in AI researchBenchmarking LLM agents for real-world dataset curationDetecting subtle data flaws using LLM agents

LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Jun 26, 2024
AB
Anna Bavaresco
🏛️ University of Amsterdam | University of Trento | University of Copenhagen | Utrecht University | ETH Zürich | Saarland University | Universidade de Lisboa | LMU Munich | University of Potsdam | Heriot-Watt University | Unbabel | MCML

This study investigates whether large language models (LLMs) can reliably replace human annotators for evaluating NLP models. Method: We introduce JUDGE-BENCH—the first large-scale, multi-task, multi-dimensional automatic evaluation benchmark with high-quality human annotations—and systematically assess the effectiveness and consistency of 11 state-of-the-art LLMs as automatic evaluators across 20 NLP tasks. Our methodology integrates human annotation quality analysis, statistical significance testing, and cross-model correlation metrics (Kendall’s τ and Spearman’s ρ). Contribution/Results: LLM-based evaluation performance is highly contingent on evaluation attributes, annotator expertise level, and text source; while LLMs approximate human judgments in certain tasks, they lack universal reliability. Human annotations remain indispensable as the gold standard for pre-validation. We publicly release JUDGE-BENCH—including all human annotations, model outputs, and evaluation scripts—to advance standardized, reproducible research on LLM-based evaluation.

Assessing validity of LLMs replacing human judges in NLP evaluationsEvaluating reproducibility of proprietary vs open-weight LLM modelsMeasuring variance in LLM performance across diverse NLP tasks

This work addresses the limitations of existing evaluation methods for assessing large language models’ ability to use external tools in complex real-world scenarios, which often suffer from oversimplified toolsets, rigid workflows, or subjective scoring. To this end, we present the first large-scale benchmark grounded in real Model Context Protocol (MCP) servers, encompassing 36 MCP services, 220 tools, and 1,000 multi-step natural language tasks that require agents to autonomously discover and orchestrate multiple tools. The evaluation employs a no-tool-name prompting strategy and a fine-grained, fact-based scoring mechanism, supported by a containerized framework and multidimensional diagnostic metrics—including tool discovery, parameterization, and error recovery. Experiments reveal that state-of-the-art models achieve pass rates exceeding 50%, with primary failure modes stemming from insufficient tool utilization and task comprehension errors. The benchmark framework, task schema, and a public subset of 500 tasks are openly released.

large language modelsModel Context Protocolmulti-step workflows

Latest Papers

What's happening recently
View more

This work addresses a critical limitation in conventional large language model (LLM) evaluation, which treats benchmark datasets as homogeneous aggregates and overlooks the heterogeneity among samples in cognitive, linguistic, and task-related attributes. The authors propose a dataset-centric meta-evaluation framework that introduces fine-grained, sample-level annotations across five dimensions: cognitive demand, language quality, task characteristics, contextual dependency, and ethical safety. For the first time, this approach enables multidimensional auditing of widely used benchmarks such as MMLU and ARC. By allowing dynamic subset composition aligned with specific evaluation objectives, the framework uncovers the diversity obscured by aggregate accuracy metrics and establishes a composable evaluation paradigm tailored to targeted capabilities—such as reasoning depth or ethical sensitivity—thereby substantially enhancing the precision and interpretability of LLM assessments.

benchmark heterogeneitydataset introspectionevaluation bias

This work addresses the high cost of machine learning benchmarking by proposing a systematic framework to efficiently select small, representative subsets of datasets while preserving model ranking stability. The study presents the first comprehensive evaluation of various dataset selection strategies—including clustering, A/D-optimal experimental designs, random baselines, and a greedy farthest-first (FAFI) approach—on rank fidelity. It derives a theoretical upper bound on Spearman rank correlation error for FAFI and integrates bootstrap aggregation to yield statistically rigorous confidence intervals for comparing strategy performance. Empirical results demonstrate that as few as five datasets suffice to achieve 0.95 rank correlation in time series classification, significantly outperforming random selection in NLP tasks, though gains are limited in recommendation systems.

benchmarkingdataset selectionmodel ranking

This work addresses the lack of systematic alignment between large language models and human reviewers in survey evaluation, as well as the absence of a multidimensional, quantifiable assessment framework. To bridge this gap, the authors introduce SurveyReview—the first benchmark specifically designed for survey reviewing—comprising 675 survey papers and 1,630 structured review reports, along with standardized data splits and evaluation protocols. Building upon Qwen3-32B with LoRA fine-tuning and external knowledge augmentation, the proposed strong baseline model, SurveyAlign, translates free-form reviews into scores and justifications across four dimensions: readability, criticality, comprehensiveness, and structure. Experimental results demonstrate that SurveyAlign significantly outperforms GPT-5.2 with prompt-based evaluation, reducing average MSE from 2.28 to 1.38 and MAE from 1.15 to 0.69 on the test set, thereby substantially improving alignment with human reviewers.

benchmarklarge language modelspeer review

Hot Scholars

WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
ZZ

Zibin Zheng

IEEE Fellow, Highly Cited Researcher, Sun Yat-sen University, China
BlockchainSmart ContractServices ComputingSoftware Reliability
IS

Ion Stoica

Professor of Computer Science, UC Berkeley
Cloud ComputingNetworkingDistributed SystemsBig Data
HD

Haodong Duan

Shanghai AI Lab | CUHK | PKU
Computer VisionVideo UnderstandingMultimodal LearningGenerative AI
LB

Lei Bai

Shanghai AI Laboratory
Foundation ModelScience IntelligenceMulti-Agent SystemAutonomous Discovery