training data curation

Designs, builds, and operates datasets and the pipelines that produce them for machine learning, including selection, collection, annotation, cleaning, automated filtering and targeting, and construction of training and evaluation sets. Analyzes and implements dataset/version management, labeling workflows and automation, dataset cataloging and productization, and evaluation-dataset design and versioning to ensure quality, reproducibility, and traceability.

trainingdatacuration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$217K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Machine learning datasets pose multifaceted risks concerning technical reliability, legal compliance, and ethical legitimacy. Method: This paper introduces the first systematic, full-lifecycle dataset governance framework, integrating critical AI theory with applied data science. It employs interdisciplinary practices—including data auditing, provenance analysis, bias detection, regulatory compliance assessment, and participatory workshops—to operationalize abstract ethical principles without reliance on specific algorithms or tools. Contribution/Results: The work establishes the first generalizable governance paradigm that enables concurrent technical, legal, and ethical reflection—bridging the gap between pragmatic guidance and humanistic critique. Its open-source guidelines have been widely adopted by educators, media organizations, and open-source communities, significantly enhancing practitioners’ awareness of latent dataset risks. Moreover, the framework has directly catalyzed the publication of accountability statements and usage constraints by multiple major public datasets.

Data Utilization EfficiencyMachine Learning EthicsResponsible AI

Existing dataset documentation tools struggle to achieve real-world adoption due to ambiguous value propositions, misalignment with practical contexts, insufficient attention to human labor costs, and a lack of systemic integration. This study addresses these challenges through a mixed-methods systematic scoping review of 59 relevant publications, combining qualitative coding with quantitative synthesis to uncover the underlying motivations driving tool design and their relationship to institutional norms. The analysis identifies four key patterns that hinder adoption and advances a responsible AI design perspective that shifts emphasis from individual accountability to institutional solutions. The work advocates embedding sustainable documentation practices within organizational workflows and cultures, offering the HCI community actionable pathways toward institutionalizing responsible data stewardship.

dataset documentationdocumentation practicesResponsible AI

DataLens: ML-Oriented Interactive Tabular Data Quality Dashboard

Jan 28, 2025
MA
Mohamed Abdelaal
🏛️ Software AG | TU Darmstadt

Existing data management tools suffer from limited automation, poor interactivity, and insufficient integration with ML workflows, compromising data quality and hindering analytical and modeling performance. To address this, we propose an interactive, ML-oriented tabular data quality dashboard that establishes an adaptive, human-in-the-loop + ML-driven data cleaning闭环. Our approach integrates data profiling, multi-strategy error detection and repair—including statistical analysis, rule-based engines, and supervised/semi-supervised models—while supporting expert rule validation and labeling. Cleaning strategies are iteratively refined using downstream model performance feedback. Furthermore, we unify DataSheets, MLflow, and Delta Lake to ensure reproducibility, traceability, and versioning of the cleaning pipeline. Experiments across multiple benchmark datasets demonstrate significant improvements: error identification rate and repair accuracy increase notably, downstream ML models achieve an average 7.2% accuracy gain, and cleaning time decreases by 40%.

Data ManagementData QualityMachine Learning Integration

FlowETL: An Autonomous Example-Driven Pipeline for Data Engineering

Jul 30, 2025
MD
Mattia Di Profio
🏛️ University of Aberdeen

Existing ETL pipelines heavily rely on manual, context-sensitive design of transformation logic, resulting in poor generalizability and low reusability. To address this, we propose an example-driven autonomous ETL framework: given user-provided target data examples, it constructs a paired-sample-based planning engine that automatically infers and synthesizes high-fidelity, context-adapted data transformation programs. Integrated with modular ETL components and runtime monitoring, the framework enables end-to-end automation for multi-format, multi-structured, and multi-scale data processing. Experiments across 14 real-world, cross-domain datasets demonstrate that our approach substantially reduces human intervention while achieving high-precision transformations (average F1 score of 0.92), strong generalization across diverse schemas and formats, and practical engineering deployability.

Automating ETL workflows to reduce human interventionDesigning context-specific transformations without manual inputStandardizing diverse datasets using example-driven approaches

DCA-Bench: A Benchmark for Dataset Curation Agents

Jun 11, 2024
BH
Benhao Huang
🏛️ Shanghai Jiao Tong University | University of Illinois Urbana-Champaign | University of Michigan

Dataset quality defects—such as missing documentation, incorrect labels, and ethical risks—are pervasive in open platforms yet resistant to detection by rule-based scripts, necessitating intelligent, automated identification methods. Method: We introduce the first LLM-agent benchmark for discovering real-world dataset quality issues, covering 221 empirically validated cases across eight platforms. It uniquely evaluates agents’ ability to autonomously detect latent defects without prior prompting. We propose an automated evaluation framework powered by GPT-4o, achieving high agreement with human experts (Cohen’s κ = 0.89), and ensure benchmark reliability via multi-source real-data sampling and expert annotation. Contribution/Results: Experiments reveal that even the state-of-the-art Curator agent detects only ~30% of defects, underscoring task difficulty. All benchmark data, code, and evaluation tools are publicly released to advance intelligent data governance.

Assessing dataset quality issues in AI researchBenchmarking LLM agents for real-world dataset curationDetecting subtle data flaws using LLM agents

Latest Papers

What's happening recently
View more

ShaRE your Data! Characterizing Datasets for LLM-based Requirements Engineering

Oct 21, 2025
QM
Quim Motger
🏛️ Universitat Politècnica de Catalunya

Public datasets in the LLM4RE (Large Language Models for Requirements Engineering) domain are fragmented, poorly documented, and lack systematic description, hindering comparability and reuse. Method: We conduct the first systematic dataset mapping study in LLM4RE, analyzing 62 publicly available datasets drawn from 43 scholarly publications along dimensions including document type, granularity, RE task phase, domain, and language. We propose the first domain-specific dataset classification and characterization framework for LLM4RE. Contribution/Results: Our framework identifies critical research gaps—particularly in requirements elicitation, requirements management, and multilingual support—and we release an open dataset catalog alongside a standardized featureization schema. This work significantly enhances dataset visibility, structural consistency, and cross-study comparability, laying the foundation for a unified benchmarking repository in LLM4RE.

Addressing limited visibility of datasets in LLM4RE researchCharacterizing fragmented datasets for LLM-based Requirements EngineeringInvestigating under-represented RE tasks and dataset descriptors

This work addresses the critical yet under-automated role of data engineering in modern machine learning systems, where performance heavily relies on high-quality data. The authors propose DataMaster, a framework that autonomously optimizes the data pipeline—encompassing external data discovery, selection, cleaning, and transformation—conditioned on the downstream learning task to enhance the performance of a fixed learning algorithm. Its key innovations include a DataTree structure to organize search branches, a shared Data Pool for reusable data assets, and a Global Memory mechanism enabling cross-branch knowledge transfer, collectively tackling the challenges of open-ended search spaces and delayed reward signals. Experiments demonstrate that DataMaster improves the medal rate by 32.27% on MLE-Bench Lite and achieves 31.02% accuracy on the GPQA task in PostTrainBench, significantly outperforming existing instruction-tuned models.

autonomous data engineeringdata discoverydata optimization

Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.

benchmarkingdataset qualitymodel datasets

Hierarchical Dataset Selection for High-Quality Data Sharing

Dec 11, 2025
XZ
Xiaona Zhou
🏛️ University of Illinois Urbana-Champaign | University of Cincinnati | Virginia Polytechnic Institute and State University

In multi-source, heterogeneous data sharing scenarios, efficiently selecting high-value datasets to enhance downstream model performance remains challenging. Method: This paper formally defines the “dataset-level selection” task and proposes a two-tier utility modeling framework that jointly captures heterogeneity across both datasets and data sources (e.g., institutions, domains), enabling few-shot generalization and adaptive selection under resource constraints. We introduce Dataset Selection via Hierarchies (DaSH), integrating hierarchical Bayesian modeling with utility propagation to jointly optimize active exploration and decision-making. Results: Experiments on Digit-Five and DomainNet demonstrate up to a 26.2% accuracy improvement over baselines, significantly reduced exploration steps, and strong robustness under low-resource conditions and critical data absence—establishing a new paradigm for principled, scalable dataset selection in heterogeneous federated settings.

Improving downstream performance under resource constraintsModeling utility at dataset and group levels for efficiencySelecting high-quality datasets from heterogeneous sources

Hot Scholars

WL

Wenxuan Li

Johns Hopkins University
Imaging InformaticsComputer-aided Diagnosis
QY

Qian Yu

Professor, Dept of Earth, Geographic, and Climate Sciences, University of Massachusetts-Amherst
GISremote sensingSpatial modeling
TZ

Tiezheng Zhang

Johns Hopkins University
Computer VisionCognitive ScienceMedical Image Analysis
YC

Yixiong Chen

Johns Hopkins University
Vision Language ModelsComputer VisionMedical Image Analysis
YC

Yining Cao

Ph.D student, University of California, San Diego
Human-Computer Interaction