dataset engineering

Designs and implements end-to-end dataset assets and their processing pipelines, including manual and expert curation, multilingual and resource curation, cross-verification filtering, and workflows for selecting, partitioning, preprocessing, and perturbing data. Builds and applies methods to distill, rank, and evaluate dataset quality and representativeness—such as response or dataset distillation, diversity-aware sample selection, difficulty grading, similarity-based ranking—and analyzes how dataset composition and handling choices affect downstream model behavior and performance.

datasetengineering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.62
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$206K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Hierarchical Dataset Selection for High-Quality Data Sharing

Dec 11, 2025
XZ
Xiaona Zhou
🏛️ University of Illinois Urbana-Champaign | University of Cincinnati | Virginia Polytechnic Institute and State University

In multi-source, heterogeneous data sharing scenarios, efficiently selecting high-value datasets to enhance downstream model performance remains challenging. Method: This paper formally defines the “dataset-level selection” task and proposes a two-tier utility modeling framework that jointly captures heterogeneity across both datasets and data sources (e.g., institutions, domains), enabling few-shot generalization and adaptive selection under resource constraints. We introduce Dataset Selection via Hierarchies (DaSH), integrating hierarchical Bayesian modeling with utility propagation to jointly optimize active exploration and decision-making. Results: Experiments on Digit-Five and DomainNet demonstrate up to a 26.2% accuracy improvement over baselines, significantly reduced exploration steps, and strong robustness under low-resource conditions and critical data absence—establishing a new paradigm for principled, scalable dataset selection in heterogeneous federated settings.

Improving downstream performance under resource constraintsModeling utility at dataset and group levels for efficiencySelecting high-quality datasets from heterogeneous sources

DataS^3: Dataset Subset Selection for Specialization

Apr 22, 2025
NH
Neha Hulkund
🏛️ MIT | University of Haifa | Woods Hole Oceanographic Institution | UC Berkeley | Princeton University | Stanford University | McGill University | Mila - Quebec AI Institute | University of Toronto | Arizona State University | UCSF

This paper addresses performance degradation of machine learning models in specific deployment environments (e.g., a hospital or national park) due to distributional shift. We formalize the Deployment-Specialized Subset Selection (DS3) task: selecting an optimal subset from a general training set to maximize model performance in the target environment. We introduce DataS³, the first cross-domain, multi-scenario benchmark for DS3, and empirically demonstrate the systematic failure of mainstream data selection methods on this task. To overcome these limitations, we propose a novel DS3 framework integrating coresets, distribution-aware filtering, and human curation—enabling subset optimization even without labeled deployment data. Experiments show that expert-curated subsets yield average accuracy gains of 20.7%, with peaks up to 51.3%, significantly outperforming full-dataset training and existing selection baselines. Moreover, DS3 improves training efficiency and enhances generalization robustness across diverse domains.

Addresses imbalanced data distributions in specialized real-world applicationsBenchmarks methods for deployment-focused dataset curation and efficiencySelects training subsets to optimize deployment-specific ML performance

Machine learning (ML) suffers from weak data curation practices and insufficient documentation of ethical, environmental, and data management information. Method: We systematically evaluated 60 datasets from the NeurIPS Datasets and Benchmarks Track (2021–2023), introducing bibliometric data cataloging theory from library and information science to ML for the first time. We developed a literature-driven, four-dimensional evaluation framework—assessing documentation completeness, ethical impact, environmental footprint, and data management—and designed an actionable, structured rubric alongside an open-source assessment toolkit. Contribution/Results: We released the first exemplar metadata repository showcasing best practices. Our analysis revealed widespread deficiencies across all four dimensions. Based on these findings, we formulated actionable guidelines for conference reviewers and community adoption. All artifacts—including framework, rubric, toolkit, and metadata—are openly shared to advance ML datasets toward higher quality, reusability, and standardization.

Data CurationEthical ConsiderationsMachine Learning

Not All Samples Should Be Utilized Equally: Towards Understanding and Improving Dataset Distillation

Aug 22, 2024
SW
Shaobo Wang
🏛️ Shanghai Jiao Tong University | Harbin Institute of Technology | National University of Singapore

Dataset distillation (DD) suffers from performance degradation under low images-per-class (IPC) regimes and lacks a theoretical understanding of how sample difficulty affects distillation. This paper presents the first unified analysis of matching-based DD methods from the perspective of sample difficulty, revealing their implicit bias toward easily learnable samples. We establish a theoretical framework for quantifying sample difficulty based on gradient norm magnitude. Building upon this insight, we propose Sample Difficulty Correction (SDC), a plug-and-play mechanism that explicitly steers the distillation process to prioritize synthesizing easily learnable samples. SDC integrates gradient-norm-based difficulty measurement, an extended neural scaling law, and optimized matching loss. Evaluated across six benchmarks and seven baseline methods, SDC consistently improves distilled dataset accuracy—yielding average gains of 2.1–5.7 percentage points under low-IPC settings. Our work establishes an interpretable, reusable, difficulty-aware paradigm for dataset distillation.

Improving distilled dataset quality by prioritizing easier samplesTheoretical explanation of matching-based methods via neural scaling lawsUnderstanding dataset distillation through sample difficulty analysis

What Makes a Good Dataset for Knowledge Distillationƒ

Nov 19, 2024
LF
Logan Frank
🏛️ Ohio State University

In knowledge distillation, the unavailability of the teacher model’s original training data—due to constraints such as continual learning or data privacy—poses a critical practical bottleneck. Method: This paper systematically investigates the efficacy of substitute datasets for data-free distillation. It proposes and validates non-natural images (e.g., StyleGAN-generated samples) as effective distillation sources, challenging the conventional assumption that original data is indispensable. A multi-dimensional evaluation framework is introduced to quantify distillation data quality along axes of diversity, discriminability, and feature alignment with the teacher. Contribution/Results: Through cross-domain data assessment, teacher–student feature alignment analysis, and ablation studies, the work demonstrates that diverse real and synthetic substitutes achieve distillation performance on par with original data on benchmarks like CIFAR-100—yielding up to a 3.2% accuracy gain in student models. This establishes a novel paradigm and practical guidelines for data-free knowledge distillation.

Establishing criteria for suitable knowledge transfer datasets in KDExploring surrogate datasets including synthetic imagery for distillationIdentifying effective datasets for knowledge distillation without original training data

Latest Papers

What's happening recently
View more

This study addresses the lack of a standardized evaluation protocol in dataset distillation research, which has hindered objective comparisons between distilled datasets and real-data baselines such as coreset methods. Under a unified experimental setup, the authors conduct the first systematic comparison of seven state-of-the-art distillation techniques against three coreset selection strategies across ImageNet-1K, ImageNet100, and ImageNette, employing both standard empirical risk minimization (ERM) and single/multi-teacher training protocols. Comprehensive evaluations along dimensions of accuracy, representativeness, diversity, and distributional coverage reveal that current distillation approaches do not consistently outperform—and often underperform—coreset methods on large-scale datasets, despite incurring substantially higher computational costs. Notably, coresets demonstrate superior coverage of the original data distribution.

coreset selectiondata efficiencydataset distillation

This work addresses the limitation of existing dataset distillation methods, which often overlook high-level semantic information and struggle to balance class discriminability with sample diversity. To overcome this, the authors propose a semantic-aware dataset distillation framework that leverages CLIP as a semantic prior for the first time. They introduce three semantic scoring functions and a two-stage sampling strategy: first selecting samples with strong semantic discriminability, then dynamically choosing diverse instances to minimize redundancy. This approach systematically integrates class relevance, inter-class separability, and intra-set diversity in the semantic space, establishing a semantic-driven criterion for efficient dataset compression. Extensive experiments across multiple datasets, image pools, and downstream models demonstrate that the proposed method consistently outperforms current state-of-the-art approaches, confirming the effectiveness and generalizability of incorporating semantic information into dataset distillation.

class-discriminativecompact datasetdataset distillation

This work addresses the high cost of machine learning benchmarking by proposing a systematic framework to efficiently select small, representative subsets of datasets while preserving model ranking stability. The study presents the first comprehensive evaluation of various dataset selection strategies—including clustering, A/D-optimal experimental designs, random baselines, and a greedy farthest-first (FAFI) approach—on rank fidelity. It derives a theoretical upper bound on Spearman rank correlation error for FAFI and integrates bootstrap aggregation to yield statistically rigorous confidence intervals for comparing strategy performance. Empirical results demonstrate that as few as five datasets suffice to achieve 0.95 rank correlation in time series classification, significantly outperforming random selection in NLP tasks, though gains are limited in recommendation systems.

benchmarkingdataset selectionmodel ranking

Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.

benchmarkingdataset qualitymodel datasets

This work addresses the limitation of existing AI benchmarks, which predominantly assess isolated data science capabilities while neglecting systematic evaluation of end-to-end project completion. The authors propose the first comprehensive evaluation framework tailored to full-cycle data science projects, introducing a benchmark comprising 40 real-world tasks that integrate multidimensional competencies—including technical implementation, analytical reasoning, communication, and ethical considerations. They further develop an assessment pipeline combining structured scoring rubrics with automated evaluation procedures. Experimental results demonstrate that state-of-the-art generative AI models perform comparably to junior data scientists on well-structured tasks, yet exhibit substantial performance gaps in tasks requiring subjective judgment, thereby underscoring the continued necessity of human validation in complex data science workflows.

AI benchmarkingautomated evaluationdata science workflow

Hot Scholars

AD

Andreas Dengel

Professor of Computer Science, University of Kaiserslautern & Executive Director, DFKI
Artificial IntelligenceMachine LearningDocument AnalysisSemantic Technologies
HS

Horst Samulowitz

IBM Research
Artificial IntelligenceAI for AICombinatorial OptimizationMeta-Learning
ZM

Zhipeng Ma

Southwest Jiaotong University
Data-Centric AILarge Language ModelHuman Mobility
BN

Bo Nørregaard Jørgensen

Professor, PhD., Head of Center for Energy Informatics, University of Southern Denmark
Energy InformaticsEnergy-ecosystemsAI AgentsMulti-agent systems
ES

Eric Sax

Karlsruhe Institute for Technology
Systems Engineering