Score
Designs and implements end-to-end dataset assets and their processing pipelines, including manual and expert curation, multilingual and resource curation, cross-verification filtering, and workflows for selecting, partitioning, preprocessing, and perturbing data. Builds and applies methods to distill, rank, and evaluate dataset quality and representativeness—such as response or dataset distillation, diversity-aware sample selection, difficulty grading, similarity-based ranking—and analyzes how dataset composition and handling choices affect downstream model behavior and performance.
Traditional data preparation methods face limitations in semantic understanding and generalization, struggling to meet the rapidly growing demand for application-ready data. This work systematically reviews the application of large language models (LLMs) in three core tasks—data cleaning, integration, and augmentation—and proposes a task-centered taxonomy that, for the first time, delineates the evolutionary trajectory of LLM-driven data preparation techniques. Through a comprehensive literature review, the study examines key technologies such as prompt engineering, agent-based architectures, and semantic matching, alongside prevailing datasets and evaluation metrics. It highlights LLMs’ strengths in enhancing generalization and semantic comprehension while identifying critical challenges related to computational cost, hallucination, scalability, and the lack of standardized evaluation frameworks. The paper concludes by outlining a roadmap for future research and development in this emerging field.
In multi-source, heterogeneous data sharing scenarios, efficiently selecting high-value datasets to enhance downstream model performance remains challenging. Method: This paper formally defines the “dataset-level selection” task and proposes a two-tier utility modeling framework that jointly captures heterogeneity across both datasets and data sources (e.g., institutions, domains), enabling few-shot generalization and adaptive selection under resource constraints. We introduce Dataset Selection via Hierarchies (DaSH), integrating hierarchical Bayesian modeling with utility propagation to jointly optimize active exploration and decision-making. Results: Experiments on Digit-Five and DomainNet demonstrate up to a 26.2% accuracy improvement over baselines, significantly reduced exploration steps, and strong robustness under low-resource conditions and critical data absence—establishing a new paradigm for principled, scalable dataset selection in heterogeneous federated settings.
This paper addresses performance degradation of machine learning models in specific deployment environments (e.g., a hospital or national park) due to distributional shift. We formalize the Deployment-Specialized Subset Selection (DS3) task: selecting an optimal subset from a general training set to maximize model performance in the target environment. We introduce DataS³, the first cross-domain, multi-scenario benchmark for DS3, and empirically demonstrate the systematic failure of mainstream data selection methods on this task. To overcome these limitations, we propose a novel DS3 framework integrating coresets, distribution-aware filtering, and human curation—enabling subset optimization even without labeled deployment data. Experiments show that expert-curated subsets yield average accuracy gains of 20.7%, with peaks up to 51.3%, significantly outperforming full-dataset training and existing selection baselines. Moreover, DS3 improves training efficiency and enhances generalization robustness across diverse domains.
Machine learning (ML) suffers from weak data curation practices and insufficient documentation of ethical, environmental, and data management information. Method: We systematically evaluated 60 datasets from the NeurIPS Datasets and Benchmarks Track (2021–2023), introducing bibliometric data cataloging theory from library and information science to ML for the first time. We developed a literature-driven, four-dimensional evaluation framework—assessing documentation completeness, ethical impact, environmental footprint, and data management—and designed an actionable, structured rubric alongside an open-source assessment toolkit. Contribution/Results: We released the first exemplar metadata repository showcasing best practices. Our analysis revealed widespread deficiencies across all four dimensions. Based on these findings, we formulated actionable guidelines for conference reviewers and community adoption. All artifacts—including framework, rubric, toolkit, and metadata—are openly shared to advance ML datasets toward higher quality, reusability, and standardization.
Dataset distillation (DD) suffers from performance degradation under low images-per-class (IPC) regimes and lacks a theoretical understanding of how sample difficulty affects distillation. This paper presents the first unified analysis of matching-based DD methods from the perspective of sample difficulty, revealing their implicit bias toward easily learnable samples. We establish a theoretical framework for quantifying sample difficulty based on gradient norm magnitude. Building upon this insight, we propose Sample Difficulty Correction (SDC), a plug-and-play mechanism that explicitly steers the distillation process to prioritize synthesizing easily learnable samples. SDC integrates gradient-norm-based difficulty measurement, an extended neural scaling law, and optimized matching loss. Evaluated across six benchmarks and seven baseline methods, SDC consistently improves distilled dataset accuracy—yielding average gains of 2.1–5.7 percentage points under low-IPC settings. Our work establishes an interpretable, reusable, difficulty-aware paradigm for dataset distillation.
In knowledge distillation, the unavailability of the teacher model’s original training data—due to constraints such as continual learning or data privacy—poses a critical practical bottleneck. Method: This paper systematically investigates the efficacy of substitute datasets for data-free distillation. It proposes and validates non-natural images (e.g., StyleGAN-generated samples) as effective distillation sources, challenging the conventional assumption that original data is indispensable. A multi-dimensional evaluation framework is introduced to quantify distillation data quality along axes of diversity, discriminability, and feature alignment with the teacher. Contribution/Results: Through cross-domain data assessment, teacher–student feature alignment analysis, and ablation studies, the work demonstrates that diverse real and synthetic substitutes achieve distillation performance on par with original data on benchmarks like CIFAR-100—yielding up to a 3.2% accuracy gain in student models. This establishes a novel paradigm and practical guidelines for data-free knowledge distillation.
This study addresses the lack of a standardized evaluation protocol in dataset distillation research, which has hindered objective comparisons between distilled datasets and real-data baselines such as coreset methods. Under a unified experimental setup, the authors conduct the first systematic comparison of seven state-of-the-art distillation techniques against three coreset selection strategies across ImageNet-1K, ImageNet100, and ImageNette, employing both standard empirical risk minimization (ERM) and single/multi-teacher training protocols. Comprehensive evaluations along dimensions of accuracy, representativeness, diversity, and distributional coverage reveal that current distillation approaches do not consistently outperform—and often underperform—coreset methods on large-scale datasets, despite incurring substantially higher computational costs. Notably, coresets demonstrate superior coverage of the original data distribution.
This work addresses the limitation of existing dataset distillation methods, which often overlook high-level semantic information and struggle to balance class discriminability with sample diversity. To overcome this, the authors propose a semantic-aware dataset distillation framework that leverages CLIP as a semantic prior for the first time. They introduce three semantic scoring functions and a two-stage sampling strategy: first selecting samples with strong semantic discriminability, then dynamically choosing diverse instances to minimize redundancy. This approach systematically integrates class relevance, inter-class separability, and intra-set diversity in the semantic space, establishing a semantic-driven criterion for efficient dataset compression. Extensive experiments across multiple datasets, image pools, and downstream models demonstrate that the proposed method consistently outperforms current state-of-the-art approaches, confirming the effectiveness and generalizability of incorporating semantic information into dataset distillation.
This work addresses the high cost of machine learning benchmarking by proposing a systematic framework to efficiently select small, representative subsets of datasets while preserving model ranking stability. The study presents the first comprehensive evaluation of various dataset selection strategies—including clustering, A/D-optimal experimental designs, random baselines, and a greedy farthest-first (FAFI) approach—on rank fidelity. It derives a theoretical upper bound on Spearman rank correlation error for FAFI and integrates bootstrap aggregation to yield statistically rigorous confidence intervals for comparing strategy performance. Empirical results demonstrate that as few as five datasets suffice to achieve 0.95 rank correlation in time series classification, significantly outperforming random selection in NLP tasks, though gains are limited in recommendation systems.
Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.
This work addresses the limitation of existing AI benchmarks, which predominantly assess isolated data science capabilities while neglecting systematic evaluation of end-to-end project completion. The authors propose the first comprehensive evaluation framework tailored to full-cycle data science projects, introducing a benchmark comprising 40 real-world tasks that integrate multidimensional competencies—including technical implementation, analytical reasoning, communication, and ethical considerations. They further develop an assessment pipeline combining structured scoring rubrics with automated evaluation procedures. Experimental results demonstrate that state-of-the-art generative AI models perform comparably to junior data scientists on well-structured tasks, yet exhibit substantial performance gaps in tasks requiring subjective judgment, thereby underscoring the continued necessity of human validation in complex data science workflows.