dataset collection and acquisition

Designs and builds datasets and the pipelines that acquire, construct, and organize data at project or large scale, specifying collection protocols, sampling and steerable selection strategies, preprocessing steps, annotation schemas, and quality-control procedures. Implements and evaluates curation workflows and tools — including automated and LLM-assisted methods — to produce benchmark, long‑tail, low‑resource, domain- or modality‑specific (e.g., audio, geospatial, multimodal) datasets and to manage dataset provenance, lifecycle, and readiness for training and evaluation.

datasetcollectionandacquisition

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.55
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$231K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Current data processing pipelines for post-training large language models—encompassing cleaning, deduplication, synthesis, and quality filtering—are fragmented and lack auditability and sample-level decision transparency. This work proposes the first end-to-end configurable data processing framework that unifies data ingestion, cleaning, LLM-driven synthesis across eight task types, three-tiered quality gating, and export modules. The system introduces sample-level provenance tracking and a precise hallucination verification mechanism. It supports six input formats and over 100 model APIs via LiteLLM, offering both a YAML-driven command-line interface and a Python API. Outputs are compatible with five training formats used by TRL, Unsloth, and AlignTune, substantially enhancing transparency, reproducibility, and scalability in post-training data preparation.

data curationLLM post-trainingpipeline auditability

ShaRE your Data! Characterizing Datasets for LLM-based Requirements Engineering

Oct 21, 2025
QM
Quim Motger
🏛️ Universitat Politècnica de Catalunya

Public datasets in the LLM4RE (Large Language Models for Requirements Engineering) domain are fragmented, poorly documented, and lack systematic description, hindering comparability and reuse. Method: We conduct the first systematic dataset mapping study in LLM4RE, analyzing 62 publicly available datasets drawn from 43 scholarly publications along dimensions including document type, granularity, RE task phase, domain, and language. We propose the first domain-specific dataset classification and characterization framework for LLM4RE. Contribution/Results: Our framework identifies critical research gaps—particularly in requirements elicitation, requirements management, and multilingual support—and we release an open dataset catalog alongside a standardized featureization schema. This work significantly enhances dataset visibility, structural consistency, and cross-study comparability, laying the foundation for a unified benchmarking repository in LLM4RE.

Addressing limited visibility of datasets in LLM4RE researchCharacterizing fragmented datasets for LLM-based Requirements EngineeringInvestigating under-represented RE tasks and dataset descriptors

Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability

Jun 02, 2025
GW
G. Winata
🏛️ Capital One | Stanford University | Carnegie Mellon University | MBZUAI | University of Toronto | ITB | Brown University | Columbia University | Oracle | Ontario Tech University | Monash University

Contemporary dataset papers frequently suffer from limited originality, insufficient diversity, inadequate quality control, and poor transparency regarding construction methodologies; existing datasheets are largely descriptive and lack quantifiable evaluation criteria or enforceable accountability mechanisms. Method: We propose DataRubrics, the first rubric-based framework for structured data quality assessment, integrating LLM-as-a-judge (e.g., GPT-4) with synthetic data techniques to enable automated, reproducible, and standardized quality scoring for both human- and model-generated datasets. Contribution/Results: The framework delivers an open-source evaluation toolkit (github.com/datarubrics/datarubrics), facilitating collaborative, measurable data review by reviewers and authors alike. It significantly enhances rigor, transparency, and trustworthiness in data-centric research through objective, interpretable, and auditable quality metrics.

Insufficient transparency in dataset construction detailsLack standardized metrics for dataset quality evaluationNeed scalable methods for synthetic data generation

Machine learning (ML) suffers from weak data curation practices and insufficient documentation of ethical, environmental, and data management information. Method: We systematically evaluated 60 datasets from the NeurIPS Datasets and Benchmarks Track (2021–2023), introducing bibliometric data cataloging theory from library and information science to ML for the first time. We developed a literature-driven, four-dimensional evaluation framework—assessing documentation completeness, ethical impact, environmental footprint, and data management—and designed an actionable, structured rubric alongside an open-source assessment toolkit. Contribution/Results: We released the first exemplar metadata repository showcasing best practices. Our analysis revealed widespread deficiencies across all four dimensions. Based on these findings, we formulated actionable guidelines for conference reviewers and community adoption. All artifacts—including framework, rubric, toolkit, and metadata—are openly shared to advance ML datasets toward higher quality, reusability, and standardization.

Data CurationEthical ConsiderationsMachine Learning

AutoDCWorkflow: LLM-based Data Cleaning Workflow Auto-Generation and Benchmark

Dec 09, 2024
LL
Lan Li
🏛️ University of Illinois, Urbana-Champaign

Data cleaning remains highly manual, inefficient, and error-prone. This paper proposes the first goal-driven LLM-based framework for automatic workflow generation: given a dirty table and a target query, it end-to-end generates a minimal viable clean table along with executable cleaning steps—including deduplication, missing-value imputation, and format standardization. Our contributions are threefold: (1) We introduce the first benchmark dataset comprising annotated quadruples of (goal, dirty table, cleaning workflow, cleaned answer); (2) We design a zero-shot, multi-stage prompting framework—requiring no fine-tuning—that decomposes the task into goal column identification, data quality diagnosis, and operation-parameter generation; (3) We empirically validate that off-the-shelf LLMs possess inherent reasoning capabilities sufficient to generate high-quality, executable cleaning workflows across three major LLM families, significantly reducing human intervention.

Addressing format inconsistencies, type errors, duplicates in datasetsAutomating data cleaning workflow generation using LLMsEvaluating workflow quality against human-curated benchmarks

Latest Papers

What's happening recently
View more

This study addresses the challenge of low-quality metadata that hinders dataset discoverability and reuse, particularly in the context of large language model (LLM)-generated descriptions lacking empirical guidance on context selection and its impact on quality. Building a literature-based framework for description quality assessment, the authors conduct systematic ablation experiments across 252 real-world CSV datasets. They uncover a previously unreported “table-structure penalty” phenomenon: relying solely on table structure significantly degrades narrative quality. While representative data samples aid semantic grounding, they do not improve overall human-rated quality. The work further reveals that different LLMs exhibit consistent descriptive styles. Through LLM-as-a-judge evaluations, semantic attribute analysis, and large-scale experimentation, the study offers key recommendations for LLM-assisted data publishing: concise, relevant context yields better results than redundant input, and table structure should be used cautiously as a basis for generation.

context ablationdata reusedataset description

This study investigates how domain-specific metadata schemas can be effectively integrated with the generic DataCite schema to enhance metadata quality and interoperability in research data repositories. Through structural comparisons, cross-schema mapping analyses, and workflow evaluations of metadata records from eight repositories in the earth and social sciences, the research reveals how disciplinary characteristics influence the completeness of DataCite records. Findings indicate that discrepancies between schemas stem primarily from differing modeling philosophies rather than expressive capacity. While optimized cross-schema mappings significantly improve metadata quality, the diversity of repository workflows also critically affects record completeness. Building on these insights, the study proposes a strategy that leverages the complementary strengths of domain-specific and generic schemas, offering practical guidance for fostering interdisciplinary data sharing.

DataCitedisciplinary metadatametadata interoperability

Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.

benchmarkingdataset qualitymodel datasets

Existing Model Cards and Data Cards describe only static models and datasets, lacking documentation of the execution context surrounding generation, transformation, and evaluation processes—thereby limiting reproducibility and bias analysis. This work proposes Workflow Cards, which extend the structured documentation paradigm to dynamic workflow executions for the first time. Built upon provenance data, Workflow Cards generate machine-readable, structured summaries interpretable by both humans and large language models (LLMs), and incorporate a template designed to answer typical execution-related questions. Experimental results demonstrate that Workflow Cards substantially enhance understanding of workflow executions compared to schema-based query interfaces, nearly doubling answer quality and achieving superior performance under both LLM-as-a-Judge and human evaluations.

Data CardsModel Cardsprovenance data

Hot Scholars

SW

Shinji Watanabe

Carnegie Mellon University
Speech recognitionSpeech processingSpeech enhancementSpeech translation
LB

Lidong Bing

MiroMind, Alibaba DAMO, Tencent, CMU, CUHK
Natural Language ProcessingLarge Language ModelsLarge Multimodal Models
GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
WS

Wenqi Shao

Researcher at Shanghai AI Laboratory
Foundation Model EvaluationLLM CompressionEfficient AdaptationMultimodal Learning