Score
Designs, builds, and operates datasets and the pipelines that produce them for machine learning, including selection, collection, annotation, cleaning, automated filtering and targeting, and construction of training and evaluation sets. Analyzes and implements dataset/version management, labeling workflows and automation, dataset cataloging and productization, and evaluation-dataset design and versioning to ensure quality, reproducibility, and traceability.
Machine learning datasets pose multifaceted risks concerning technical reliability, legal compliance, and ethical legitimacy. Method: This paper introduces the first systematic, full-lifecycle dataset governance framework, integrating critical AI theory with applied data science. It employs interdisciplinary practices—including data auditing, provenance analysis, bias detection, regulatory compliance assessment, and participatory workshops—to operationalize abstract ethical principles without reliance on specific algorithms or tools. Contribution/Results: The work establishes the first generalizable governance paradigm that enables concurrent technical, legal, and ethical reflection—bridging the gap between pragmatic guidance and humanistic critique. Its open-source guidelines have been widely adopted by educators, media organizations, and open-source communities, significantly enhancing practitioners’ awareness of latent dataset risks. Moreover, the framework has directly catalyzed the publication of accountability statements and usage constraints by multiple major public datasets.
Existing dataset documentation tools struggle to achieve real-world adoption due to ambiguous value propositions, misalignment with practical contexts, insufficient attention to human labor costs, and a lack of systemic integration. This study addresses these challenges through a mixed-methods systematic scoping review of 59 relevant publications, combining qualitative coding with quantitative synthesis to uncover the underlying motivations driving tool design and their relationship to institutional norms. The analysis identifies four key patterns that hinder adoption and advances a responsible AI design perspective that shifts emphasis from individual accountability to institutional solutions. The work advocates embedding sustainable documentation practices within organizational workflows and cultures, offering the HCI community actionable pathways toward institutionalizing responsible data stewardship.
Existing data management tools suffer from limited automation, poor interactivity, and insufficient integration with ML workflows, compromising data quality and hindering analytical and modeling performance. To address this, we propose an interactive, ML-oriented tabular data quality dashboard that establishes an adaptive, human-in-the-loop + ML-driven data cleaning闭环. Our approach integrates data profiling, multi-strategy error detection and repair—including statistical analysis, rule-based engines, and supervised/semi-supervised models—while supporting expert rule validation and labeling. Cleaning strategies are iteratively refined using downstream model performance feedback. Furthermore, we unify DataSheets, MLflow, and Delta Lake to ensure reproducibility, traceability, and versioning of the cleaning pipeline. Experiments across multiple benchmark datasets demonstrate significant improvements: error identification rate and repair accuracy increase notably, downstream ML models achieve an average 7.2% accuracy gain, and cleaning time decreases by 40%.
Existing ETL pipelines heavily rely on manual, context-sensitive design of transformation logic, resulting in poor generalizability and low reusability. To address this, we propose an example-driven autonomous ETL framework: given user-provided target data examples, it constructs a paired-sample-based planning engine that automatically infers and synthesizes high-fidelity, context-adapted data transformation programs. Integrated with modular ETL components and runtime monitoring, the framework enables end-to-end automation for multi-format, multi-structured, and multi-scale data processing. Experiments across 14 real-world, cross-domain datasets demonstrate that our approach substantially reduces human intervention while achieving high-precision transformations (average F1 score of 0.92), strong generalization across diverse schemas and formats, and practical engineering deployability.
Dataset quality defects—such as missing documentation, incorrect labels, and ethical risks—are pervasive in open platforms yet resistant to detection by rule-based scripts, necessitating intelligent, automated identification methods. Method: We introduce the first LLM-agent benchmark for discovering real-world dataset quality issues, covering 221 empirically validated cases across eight platforms. It uniquely evaluates agents’ ability to autonomously detect latent defects without prior prompting. We propose an automated evaluation framework powered by GPT-4o, achieving high agreement with human experts (Cohen’s κ = 0.89), and ensure benchmark reliability via multi-source real-data sampling and expert annotation. Contribution/Results: Experiments reveal that even the state-of-the-art Curator agent detects only ~30% of defects, underscoring task difficulty. All benchmark data, code, and evaluation tools are publicly released to advance intelligent data governance.
Public datasets in the LLM4RE (Large Language Models for Requirements Engineering) domain are fragmented, poorly documented, and lack systematic description, hindering comparability and reuse. Method: We conduct the first systematic dataset mapping study in LLM4RE, analyzing 62 publicly available datasets drawn from 43 scholarly publications along dimensions including document type, granularity, RE task phase, domain, and language. We propose the first domain-specific dataset classification and characterization framework for LLM4RE. Contribution/Results: Our framework identifies critical research gaps—particularly in requirements elicitation, requirements management, and multilingual support—and we release an open dataset catalog alongside a standardized featureization schema. This work significantly enhances dataset visibility, structural consistency, and cross-study comparability, laying the foundation for a unified benchmarking repository in LLM4RE.
This work addresses the critical yet under-automated role of data engineering in modern machine learning systems, where performance heavily relies on high-quality data. The authors propose DataMaster, a framework that autonomously optimizes the data pipeline—encompassing external data discovery, selection, cleaning, and transformation—conditioned on the downstream learning task to enhance the performance of a fixed learning algorithm. Its key innovations include a DataTree structure to organize search branches, a shared Data Pool for reusable data assets, and a Global Memory mechanism enabling cross-branch knowledge transfer, collectively tackling the challenges of open-ended search spaces and delayed reward signals. Experiments demonstrate that DataMaster improves the medal rate by 32.27% on MLE-Bench Lite and achieves 31.02% accuracy on the GPQA task in PostTrainBench, significantly outperforming existing instruction-tuned models.
Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.
In multi-source, heterogeneous data sharing scenarios, efficiently selecting high-value datasets to enhance downstream model performance remains challenging. Method: This paper formally defines the “dataset-level selection” task and proposes a two-tier utility modeling framework that jointly captures heterogeneity across both datasets and data sources (e.g., institutions, domains), enabling few-shot generalization and adaptive selection under resource constraints. We introduce Dataset Selection via Hierarchies (DaSH), integrating hierarchical Bayesian modeling with utility propagation to jointly optimize active exploration and decision-making. Results: Experiments on Digit-Five and DomainNet demonstrate up to a 26.2% accuracy improvement over baselines, significantly reduced exploration steps, and strong robustness under low-resource conditions and critical data absence—establishing a new paradigm for principled, scalable dataset selection in heterogeneous federated settings.