evaluate data augmentation strategies

Design and run reproducible benchmarking pipelines and evaluation protocols that compare and measure the effectiveness of data augmentation techniques; build standardized metrics, controlled ablations, and reporting templates to quantify how different augmentation variants affect model performance, robustness, and generalization across datasets and backbone architectures.

evaluatedataaugmentationstrategies

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.

benchmarkingdataset qualitymodel datasets

This work identifies an evaluation bias introduced by data augmentation (e.g., SMOTE, mutation-based augmentation) in scarce-data scenarios—particularly flaky test classification—where augmented samples inadvertently contaminate the test set, severely compromising fairness and reliability assessments. To address this, the authors first empirically identify and validate the critical phenomenon that “augmented data participation in testing” induces systematic evaluation distortion. They then propose a detection framework capable of disentangling training-induced bias from evaluation-induced bias, and design a bias-calibrated evaluation protocol. Experiments across multiple flaky-test benchmark datasets demonstrate that test sets containing augmented samples inflate accuracy by up to 23.7% and introduce F1-score deviations exceeding 0.15. This study establishes both theoretical foundations and practical guidelines for trustworthy model evaluation under data augmentation.

Bias in training and testing with augmented dataEvaluating augmented data effects in model testingImpact of data augmentation on model bias

This work addresses the limitation of existing AI benchmarks, which predominantly assess isolated data science capabilities while neglecting systematic evaluation of end-to-end project completion. The authors propose the first comprehensive evaluation framework tailored to full-cycle data science projects, introducing a benchmark comprising 40 real-world tasks that integrate multidimensional competencies—including technical implementation, analytical reasoning, communication, and ethical considerations. They further develop an assessment pipeline combining structured scoring rubrics with automated evaluation procedures. Experimental results demonstrate that state-of-the-art generative AI models perform comparably to junior data scientists on well-structured tasks, yet exhibit substantial performance gaps in tasks requiring subjective judgment, thereby underscoring the continued necessity of human validation in complex data science workflows.

AI benchmarkingautomated evaluationdata science workflow

Current evaluations of AI models lack standardized protocols, with institutions selectively employing benchmarks in ways that hinder cross-study comparability and raise concerns about scientific validity. This work introduces Benchmarking-Cultures-25, a dataset encompassing 231 benchmarks from 139 model releases, and combines qualitative content analysis with a unified categorization framework to systematically expose the fragmentation in benchmark selection: 63.2% of benchmarks are used by only a single institution, and 38.5% appear just once. Moreover, many benchmarks marketed as “general-purpose” disproportionately emphasize STEM—particularly mathematics—while often neglecting construct validity. The study further proposes a taxonomy aligning ostensibly disparate terminologies to their underlying measurement signals and develops an interactive tool revealing that benchmarks frequently serve marketing narratives rather than rigorous scientific assessment.

AI evaluationbenchmarkingconstruct validity

AI Benchmark Democratization and Carpentry

Dec 12, 2025
GV
Gregor von Laszewski

Current AI benchmarking faces three critical challenges: static benchmarks are increasingly susceptible to memorization by large models, lack alignment with real-world deployment scenarios, and struggle to keep pace with rapidly evolving models, datasets, and hardware. To address these, this paper introduces a Dynamic Adaptive AI Benchmarking Framework and pioneers the “AI Benchmark Carpentry” pedagogical paradigm—integrating continuous evaluation, interpretable design, and cross-layer capability development. The framework employs a modular pipeline, dynamically refreshed datasets, lightweight cross-platform evaluation containers, and transparent provenance tracking—drawing on best practices from MLCommons and TPC. It significantly lowers the barrier to benchmark construction, enabling small- and medium-scale organizations to perform domain-specific evaluations under constrained computational resources. By shifting benchmarks from static performance leaderboards toward application-oriented decision-support tools, the framework enhances representational fidelity for deployment risks and improves result reproducibility.

Current benchmarks lack adaptability to evolving models, data, and diverse deployment scenarios.High resource demands and limited expertise hinder accessible and relevant AI evaluation.Static AI benchmarks fail to reflect real-world performance due to model memorization.

Latest Papers

What's happening recently
View more

Traditional benchmarks provide only aggregate scores, offering insufficient evidence to support reliable deployment decisions and thereby creating a disconnect between evaluation and action. To address this gap, this work proposes a “deployment-completeness” benchmarking framework, introducing novel metrics—evidence fibers, completeness curves, and certifiable proportions—alongside a systematic audit methodology comprising evidence fiber analysis, response ranking intervals, conformal coverage evaluation, and a certify-then-acquire decision pipeline. Empirical evaluation on benchmarks such as Tox21, Matbench, and JARVIS reveals that conventional approaches suffer a drastic drop in channel coverage to 10.07% under real-world deployment conditions. In contrast, the proposed method reduces error-driven deployment decisions to 0.027% on Tox21 and 0.128% on JARVIS, substantially enhancing deployment reliability.

benchmark evidencecertifiable fractiondeployment action

This study addresses a critical gap in existing tool-calling evaluation benchmarks: the lack of validation of the evaluators themselves, which risks conflating assessment artifacts with agents’ true capabilities. Through a systematic audit of four prominent benchmarks—BFCL v4, τ2-Bench, LiveMCPBench, and MCP-Atlas—the authors conduct expert review of 496 tasks, replicate experiments, and perform trajectory-level analysis, revealing an 18.5% disagreement rate between automated evaluators and human judgment. Notably, LiveMCPBench exhibits a score variance of up to 18.9 percentage points upon re-evaluation, sufficient to overturn leaderboard rankings. To address these issues, the work introduces the first unified taxonomy of tool-calling evaluation failures, advocates for distinct measurement of tool invocation, task completion, and result verification, and releases Tool-Veritas—a configurable benchmark—and Harness Lab, an open-source evaluation platform.

benchmark validityevaluator alignmentLLM benchmarks

This work addresses the limitations of existing reproducibility assessment methods, which rely on manual annotations and thus lack scalability and authentic supervision signals reflecting real-world reproduction challenges. The authors propose the first scalable evaluation framework that leverages GitHub user-submitted issues as natural supervision, enabling large-scale assessment of large language model (LLM) agents’ ability to identify paper-to-code reproducibility issues without human annotation. By integrating language understanding with code context analysis, the approach enables non-execution-based detection of reproducibility barriers. Experimental results demonstrate that the best-performing LLM agent identifies at least one semantically relevant reproducibility issue—aligned with those reported by humans—in approximately 90% of the evaluated papers, exhibiting strong performance in both failure detection and semantic localization.

benchmarkingGitHub issuesLLM agents

Existing approaches struggle to effectively quantify the similarity and quality between synthetic and real data in evaluating tool-augmented agents. To address this gap, this work proposes SynAE, a novel framework that establishes the first multi-axis evaluation system tailored for multi-turn tool-use scenarios. SynAE introduces four fine-grained metric categories—assessing task instructions, tool invocations, final outputs, and downstream evaluation performance—to systematically measure synthetic data across dimensions of validity, fidelity, and diversity. Integrating natural language processing, trajectory modeling, and controllable generation techniques, the framework enables a reproducible evaluation pipeline and successfully identifies several representative failure modes in synthetic data generation. Empirical results demonstrate that such multidimensional assessment is essential for enhancing the reliability of agent evaluations.

benchmarkingdata qualityevaluation framework

This study addresses a critical gap in current AI evaluation methodologies, which often overlook the impact of low-resource deployment conditions—such as noisy inputs, limited hardware capabilities, and unstable network connectivity—on system usability. The work proposes a novel evaluation framework that treats the deployed system as the unit of assessment, integrating task performance with real-world deployment contexts across multiple dimensions. Departing from conventional leaderboard-based approaches, the framework tailors evaluation criteria to specific application categories and introduces a standardized reporting system comprising benchmark cards, deployment profiles, and failure-handling mechanisms. By balancing comparability with contextual sensitivity, this approach provides policymakers and practitioners with clear, actionable insights for informed AI deployment decisions.

AI evaluationbenchmarkingdeployment conditions

Hot Scholars

AG

Arjun Guha

Northeastern University
Programming Languages
YH

Yang He

A*STAR & NUS
Machine LearningComputer Vision
ML

Ming Liu

Senior Lecturer of Machine Learning, Deakin University
Natural Language ProcessingMachine LearningHuman-centered AI
CZ

Chenghao Zhang

Renmin University of China
Natural Language ProcessingInformation RetrievalMultimodal
SD

Soumyaratna Debnath

School of Electrical and Electronic Engineering (EEE)
Computer VisionComputer Graphics3D Shape AnalysisPose Estimation