benchmark dataset creation

Designs and produces benchmark datasets and their evaluation artifacts by defining tasks and metrics, collecting and curating raw and labeled data, constructing standardized train/validation/test splits, and packaging metadata and releases for reuse and shared tasks. Audits and analyzes dataset quality and biases, standardizes evaluation protocols across tasks or aggregated datasets, and prepares documentation and machine‑readable artifacts that support reproducible benchmarking.

benchmarkdatasetcreation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.56
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.

benchmarkingdataset qualitymodel datasets

Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability

Jun 02, 2025
GW
G. Winata
🏛️ Capital One | Stanford University | Carnegie Mellon University | MBZUAI | University of Toronto | ITB | Brown University | Columbia University | Oracle | Ontario Tech University | Monash University

Contemporary dataset papers frequently suffer from limited originality, insufficient diversity, inadequate quality control, and poor transparency regarding construction methodologies; existing datasheets are largely descriptive and lack quantifiable evaluation criteria or enforceable accountability mechanisms. Method: We propose DataRubrics, the first rubric-based framework for structured data quality assessment, integrating LLM-as-a-judge (e.g., GPT-4) with synthetic data techniques to enable automated, reproducible, and standardized quality scoring for both human- and model-generated datasets. Contribution/Results: The framework delivers an open-source evaluation toolkit (github.com/datarubrics/datarubrics), facilitating collaborative, measurable data review by reviewers and authors alike. It significantly enhances rigor, transparency, and trustworthiness in data-centric research through objective, interpretable, and auditable quality metrics.

Insufficient transparency in dataset construction detailsLack standardized metrics for dataset quality evaluationNeed scalable methods for synthetic data generation

This work addresses a critical limitation in conventional large language model (LLM) evaluation, which treats benchmark datasets as homogeneous aggregates and overlooks the heterogeneity among samples in cognitive, linguistic, and task-related attributes. The authors propose a dataset-centric meta-evaluation framework that introduces fine-grained, sample-level annotations across five dimensions: cognitive demand, language quality, task characteristics, contextual dependency, and ethical safety. For the first time, this approach enables multidimensional auditing of widely used benchmarks such as MMLU and ARC. By allowing dynamic subset composition aligned with specific evaluation objectives, the framework uncovers the diversity obscured by aggregate accuracy metrics and establishes a composable evaluation paradigm tailored to targeted capabilities—such as reasoning depth or ethical sensitivity—thereby substantially enhancing the precision and interpretability of LLM assessments.

benchmark heterogeneitydataset introspectionevaluation bias

Existing dataset documentation tools struggle to achieve real-world adoption due to ambiguous value propositions, misalignment with practical contexts, insufficient attention to human labor costs, and a lack of systemic integration. This study addresses these challenges through a mixed-methods systematic scoping review of 59 relevant publications, combining qualitative coding with quantitative synthesis to uncover the underlying motivations driving tool design and their relationship to institutional norms. The analysis identifies four key patterns that hinder adoption and advances a responsible AI design perspective that shifts emphasis from individual accountability to institutional solutions. The work advocates embedding sustainable documentation practices within organizational workflows and cultures, offering the HCI community actionable pathways toward institutionalizing responsible data stewardship.

dataset documentationdocumentation practicesResponsible AI

DCA-Bench: A Benchmark for Dataset Curation Agents

Jun 11, 2024
BH
Benhao Huang
🏛️ Shanghai Jiao Tong University | University of Illinois Urbana-Champaign | University of Michigan

Dataset quality defects—such as missing documentation, incorrect labels, and ethical risks—are pervasive in open platforms yet resistant to detection by rule-based scripts, necessitating intelligent, automated identification methods. Method: We introduce the first LLM-agent benchmark for discovering real-world dataset quality issues, covering 221 empirically validated cases across eight platforms. It uniquely evaluates agents’ ability to autonomously detect latent defects without prior prompting. We propose an automated evaluation framework powered by GPT-4o, achieving high agreement with human experts (Cohen’s κ = 0.89), and ensure benchmark reliability via multi-source real-data sampling and expert annotation. Contribution/Results: Experiments reveal that even the state-of-the-art Curator agent detects only ~30% of defects, underscoring task difficulty. All benchmark data, code, and evaluation tools are publicly released to advance intelligent data governance.

Assessing dataset quality issues in AI researchBenchmarking LLM agents for real-world dataset curationDetecting subtle data flaws using LLM agents

Latest Papers

What's happening recently
View more

Existing benchmarks for knowledge work evaluation largely adhere to traditional NLP task paradigms, failing to capture systems’ capabilities in real-world knowledge-intensive settings. This work proposes a three-step framework—explicitly defining work activities, establishing realistic test environments, and focusing evaluation on deliverable outputs—and derives 18 core knowledge work activities from the O*NET database. Innovatively integrating role responsibilities, local tool usage, and downstream usability into benchmark design, the approach establishes a coherent “work activity–test setup–scoring artifact” alignment. Validation through three case studies (GDPval, OfficeQA Pro, and APEX-SWE) exposes critical misalignments in current benchmarks between tasks, environments, and actual work objectives, offering a new paradigm for evaluating knowledge work systems in practical, application-oriented contexts.

benchmark designevaluationknowledge work

This study addresses the limited reliability of existing training data contamination detection methods in real-world auditing scenarios, particularly when distribution shifts occur or when reference benchmarks are substantially smaller than the pretraining corpus. Through a systematic evaluation of three dominant paradigms—LLM Dataset Inference, Post-Hoc Dataset Inference, and CoDeC—the authors conduct 335 experiments across 27 open-source and state-of-the-art closed-source language models (up to 27B parameters). They identify distribution shift and small-scale benchmarks as two critical failure modes, revealing that only 199 evaluations yield correct conclusions. Current approaches suffer from high false-positive rates, low statistical power, or coarse-grained provenance resolution, rendering them inadequate for reliably verifying individual benchmark subsets and underscoring the irreplaceable value of transparent data provenance.

benchmark contaminationdata provenancedistribution shift

This study addresses a critical limitation in existing document layout analysis methods, which treat figures and tables as generic objects and thus fail to identify semantically valuable, reusable analytical visual content—referred to as “data snapshots”—in institutional documents. The work introduces the novel task of data snapshot extraction, presents a benchmark dataset comprising humanitarian reports and World Bank policy papers, and proposes an evaluation framework that integrates spatial localization with semantic annotation. Systematic evaluation of multiple open-source layout models reveals consistent shortcomings in handling institutional documents, including confusion between analytical and non-analytical content, fragmentation of composite charts, and lack of contextual awareness. By exposing the generalization bottlenecks of current models in operational documents, this research provides a foundation for future advancements through the public release of its dataset and codebase.

data snapshot extractiondocument layout analysisinstitutional documents

This study addresses a critical gap in existing tool-calling evaluation benchmarks: the lack of validation of the evaluators themselves, which risks conflating assessment artifacts with agents’ true capabilities. Through a systematic audit of four prominent benchmarks—BFCL v4, τ2-Bench, LiveMCPBench, and MCP-Atlas—the authors conduct expert review of 496 tasks, replicate experiments, and perform trajectory-level analysis, revealing an 18.5% disagreement rate between automated evaluators and human judgment. Notably, LiveMCPBench exhibits a score variance of up to 18.9 percentage points upon re-evaluation, sufficient to overturn leaderboard rankings. To address these issues, the work introduces the first unified taxonomy of tool-calling evaluation failures, advocates for distinct measurement of tool invocation, task completion, and result verification, and releases Tool-Veritas—a configurable benchmark—and Harness Lab, an open-source evaluation platform.

benchmark validityevaluator alignmentLLM benchmarks

Hot Scholars

AC

Arman Cohan

Yale University; Allen Institute for AI
Natural Language ProcessingMachine LearningArtificial Intelligence
XH

Xuming Hu

Assistant Professor, HKUST(GZ) / HKUST
Natural Language ProcessingLarge Language Model
TP

Tomas Pfister

Head of AI Research @ Google Cloud
Machine learningDeep learningComputer vision
ZL

Ziwei Liu

Associate Professor, Nanyang Technological University
Computer VisionMachine LearningComputer Graphics
PS

Philip S. Yu

Professor of Computer Science, University of Illinons at Chicago
Data miningDatabasePrivacy