build verifiable cot benchmarks

Designs and constructs evaluation benchmarks and datasets that pair tasks with attached gold chain-of-thought reasoning and executable, parameterized examples so that model outputs can be automatically checked for correctness of intermediate steps. Builds evaluation pipelines and dataset splits that support long‑horizon and evidence‑grounded problems, automated verification of numeric and logical intermediates, and contamination‑free, reproducible scoring.

buildverifiablecotbenchmarks

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses critical reliability issues in existing Lean theorem-proving benchmarks, where inconsistencies between formal statements and informal problem descriptions, along with susceptibility to trivial or adversarial solutions, undermine evaluation validity. To tackle this, the study introduces the first fault taxonomy for formal mathematical datasets and develops an automated auditing toolkit integrating static program analysis, formal verification, semantic auditing via prompt engineering, and manual validation. Applying this framework, the authors conduct a large-scale audit of five prominent benchmarks, uncovering 4,833 issues—including 398 severe defects—and demonstrate that uncorrected flaws significantly distort prover rankings. The paper releases both the auditing tools and corrected dataset snapshots to foster reproducible and trustworthy evaluations in theorem proving.

dataset defectsevaluation reliabilityformal verification

Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.

benchmarkingdataset qualitymodel datasets

This work addresses the limitation of existing AI benchmarks, which predominantly assess isolated data science capabilities while neglecting systematic evaluation of end-to-end project completion. The authors propose the first comprehensive evaluation framework tailored to full-cycle data science projects, introducing a benchmark comprising 40 real-world tasks that integrate multidimensional competencies—including technical implementation, analytical reasoning, communication, and ethical considerations. They further develop an assessment pipeline combining structured scoring rubrics with automated evaluation procedures. Experimental results demonstrate that state-of-the-art generative AI models perform comparably to junior data scientists on well-structured tasks, yet exhibit substantial performance gaps in tasks requiring subjective judgment, thereby underscoring the continued necessity of human validation in complex data science workflows.

AI benchmarkingautomated evaluationdata science workflow

BENCHAGENTS: Automated Benchmark Creation with Agent Interaction

Oct 29, 2024
NB
Natasha Butt
🏛️ University of Amsterdam | Microsoft Research | UIUC

Existing evaluation of generative AI is hindered by the scarcity of high-quality benchmarks, whose manual construction is costly and time-consuming. Method: We propose the first automated benchmark construction framework powered by collaborative large language model (LLM) agents, decomposing benchmark creation into four sequential stages—planning, generation, verification, and evaluation—integrating task decomposition, agent coordination, human-in-the-loop feedback, and explicit constraint-satisfaction assessment. Contribution/Results: The framework significantly enhances data diversity and metric reliability. Leveraging it, we construct the first high-quality benchmark specifically targeting planning and constraint-satisfaction capabilities in text generation. We systematically evaluate seven state-of-the-art models, uncovering shared failure modes and fine-grained capability disparities. Our work establishes a scalable, reproducible paradigm for evaluating generative AI capabilities, advancing both benchmark methodology and empirical analysis.

Automating high-quality benchmark creation for evolving AI modelsGenerating structured benchmarks for complex reasoning and multimodal evaluationOvercoming slow manual benchmark creation via multi-agent framework

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities

Dec 09, 2024
AG
Adhiraj Ghosh
🏛️ University of Tübingen | Open-Ψ (Open-Sci) Collective | University of Cambridge

Traditional static test sets inadequately evaluate foundation models’ diverse capabilities in open-ended scenarios. To address this, we propose ONEBench—a dynamic, extensible benchmarking paradigm that enables on-demand generation of customized evaluation suites targeting open capabilities, framing model assessment as a collective selection and aggregation process over sample-level tests. Our key contributions include: (1) the first unified, open-ended, and evolvable evaluation framework operating at the sample level; (2) a sparse measurement aggregation algorithm, a progressive sample pool construction mechanism, and a cross-modal unified interface (ONEBench-LLM/LMM); and (3) a robustness-aware scoring model with theoretical guarantees on identifiability and fast convergence. Experiments show that ONEBench achieves ranking stability >0.98 under 95% measurement sparsity, reduces evaluation cost by 20×, and attains >0.98 correlation with mean-score rankings on homogeneous data—enabling unified, efficient, and reliable assessment of both language and multimodal models.

Aggregating diverse metrics into reliable model scoresEvaluating open-ended capabilities of foundation modelsReducing evaluation cost while maintaining accuracy

Latest Papers

What's happening recently
View more

This study addresses a critical gap in existing tool-calling evaluation benchmarks: the lack of validation of the evaluators themselves, which risks conflating assessment artifacts with agents’ true capabilities. Through a systematic audit of four prominent benchmarks—BFCL v4, τ2-Bench, LiveMCPBench, and MCP-Atlas—the authors conduct expert review of 496 tasks, replicate experiments, and perform trajectory-level analysis, revealing an 18.5% disagreement rate between automated evaluators and human judgment. Notably, LiveMCPBench exhibits a score variance of up to 18.9 percentage points upon re-evaluation, sufficient to overturn leaderboard rankings. To address these issues, the work introduces the first unified taxonomy of tool-calling evaluation failures, advocates for distinct measurement of tool invocation, task completion, and result verification, and releases Tool-Veritas—a configurable benchmark—and Harness Lab, an open-source evaluation platform.

benchmark validityevaluator alignmentLLM benchmarks

This work addresses the challenge of systematically evaluating concept bottleneck models, whose applicability and failure mechanisms remain poorly understood due to the scarcity of real-world datasets with annotated concept labels. To bridge this gap, we introduce the first controllable synthetic benchmark that leverages parametric generation techniques to precisely modulate data modality, concept selection, annotation quality, and label completeness, thereby simulating diverse real-world relationships between concepts and predictions. This benchmark enables comprehensive evaluation of various concept bottleneck models across both decision-support and fully automated tasks, effectively identifying key performance determinants and characteristic failure modes. Our framework fills a critical void in the current evaluation landscape for concept-based interpretability methods.

concept bottleneck modelsconcept labelsmodel interpretability

This work addresses a critical limitation in conventional large language model (LLM) evaluation, which treats benchmark datasets as homogeneous aggregates and overlooks the heterogeneity among samples in cognitive, linguistic, and task-related attributes. The authors propose a dataset-centric meta-evaluation framework that introduces fine-grained, sample-level annotations across five dimensions: cognitive demand, language quality, task characteristics, contextual dependency, and ethical safety. For the first time, this approach enables multidimensional auditing of widely used benchmarks such as MMLU and ARC. By allowing dynamic subset composition aligned with specific evaluation objectives, the framework uncovers the diversity obscured by aggregate accuracy metrics and establishes a composable evaluation paradigm tailored to targeted capabilities—such as reasoning depth or ethical sensitivity—thereby substantially enhancing the precision and interpretability of LLM assessments.

benchmark heterogeneitydataset introspectionevaluation bias

This work addresses the lack of a unified evaluation framework for knowledge graph integration pipelines, which hinders systematic comparison and selection of methods. To bridge this gap, the paper introduces KGI-Bench, the first comprehensive benchmark specifically designed for evaluating knowledge graph data integration. KGI-Bench assesses integration performance across three key dimensions—coverage, correctness, and consistency—when incorporating heterogeneous input data (structured, semi-structured, and unstructured) into a target knowledge graph. Using a curated dataset in the movie domain, the benchmark evaluates twelve representative integration pipelines, revealing significant performance variations attributable to input data types and architectural choices. The results demonstrate the effectiveness and practical utility of KGI-Bench in enabling rigorous, reproducible evaluation of knowledge graph integration approaches.

benchmarkdata integrationknowledge graph

Existing approaches struggle to effectively quantify the similarity and quality between synthetic and real data in evaluating tool-augmented agents. To address this gap, this work proposes SynAE, a novel framework that establishes the first multi-axis evaluation system tailored for multi-turn tool-use scenarios. SynAE introduces four fine-grained metric categories—assessing task instructions, tool invocations, final outputs, and downstream evaluation performance—to systematically measure synthetic data across dimensions of validity, fidelity, and diversity. Integrating natural language processing, trajectory modeling, and controllable generation techniques, the framework enables a reproducible evaluation pipeline and successfully identifies several representative failure modes in synthetic data generation. Empirical results demonstrate that such multidimensional assessment is essential for enhancing the reliability of agent evaluations.

benchmarkingdata qualityevaluation framework

Hot Scholars

ZX

Zhifei Xie

Tsinghua University
Artificial IntelligenceLarge Multimodal ModelGPT4o
KG

Kartik Garg

Georgia Institute of Technology
Reinforcement learningRoboticsComputer Vision
SW

Siyin Wang

Fudan University
Natural Language Processing
YS

Yulei Sui

University of New South Wales (UNSW Sydney)
Static Program AnalysisSecure Software EngineeringAI4SESE4AI
JZ

Jiaheng Zhang

Assistant Professor, National University of Singapore.
Zero-knowledge proofsAI safetyApplied cryptographyBlockchain