data curation

Designs, builds, and operates processes, pipelines, and repositories for collecting, cleaning, annotating, validating, versioning, and distributing datasets. Establishes and applies metadata schemas, quality-control checks, provenance tracking, and documentation to ensure dataset usability, reproducibility, and long-term maintenance.

datacuration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$223K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing dataset documentation tools struggle to achieve real-world adoption due to ambiguous value propositions, misalignment with practical contexts, insufficient attention to human labor costs, and a lack of systemic integration. This study addresses these challenges through a mixed-methods systematic scoping review of 59 relevant publications, combining qualitative coding with quantitative synthesis to uncover the underlying motivations driving tool design and their relationship to institutional norms. The analysis identifies four key patterns that hinder adoption and advances a responsible AI design perspective that shifts emphasis from individual accountability to institutional solutions. The work advocates embedding sustainable documentation practices within organizational workflows and cultures, offering the HCI community actionable pathways toward institutionalizing responsible data stewardship.

dataset documentationdocumentation practicesResponsible AI

To address the challenges of parallel scheduling, opaque execution states, poor result reproducibility, and inadequate auditability when managing hundreds to thousands of Snakemake/Nextflow pipelines in large-scale bioinformatics analyses, this paper proposes a lightweight command-line orchestration framework. Built in Python and integrated with SQLite or PostgreSQL, it enables unified pipeline launching across heterogeneous workflows, real-time status monitoring, fine-grained log collection, automated result ingestion into databases, and comprehensive lifecycle metric logging—including runtime, resource consumption, and failure points. It introduces a novel CLI paradigm that supports cross-pipeline collaborative monitoring and reproducibility assurance without modifying existing workflow code. Experimental evaluation demonstrates a 42% improvement in multi-project throughput, significantly enhancing observability, auditability, and reproducibility in large-scale bioinformatics analysis.

Automate provisioning and evaluation of bioinformatics pipelinesCoordinate bulk processing of multiple datasets efficientlyMonitor and record pipeline metrics for reproducibility

Automatic Metadata Capture and Processing for High-Performance Workflows

Jun 18, 2025
PS
Polina Shpilker
🏛️ Tufts University | Sandia National Laboratories

In heterogeneous high-performance computing (HPC) environments, workflow metadata collection remains challenging due to fragmentation, poor reusability, and lack of standardization—hindering FAIR (Findable, Accessible, Interoperable, Reusable) compliance and impeding efficient performance analysis. To address this, we propose an automated metadata management framework. It features a lightweight runtime collection mechanism supporting fine-grained metadata capture across heterogeneous workflow systems (e.g., Snakemake, Nextflow); a dual-format unified storage scheme combining JSON Schema (for semantic expressiveness and human readability) and SQLite (for efficient querying); and a performance-aware metadata schema redesign that natively models task dependencies, resource consumption, and temporal behavior. Experimental evaluation demonstrates significant improvements in metadata findability, interoperability, and reusability—enabling reproducible, data-driven workflow performance analysis.

Apply FAIR principles to enhance research reproducibilityCapture metadata for workflows on heterogeneous architecturesStandardize and reorganize metadata for performance analysis

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

This study investigates how domain-specific metadata schemas can be effectively integrated with the generic DataCite schema to enhance metadata quality and interoperability in research data repositories. Through structural comparisons, cross-schema mapping analyses, and workflow evaluations of metadata records from eight repositories in the earth and social sciences, the research reveals how disciplinary characteristics influence the completeness of DataCite records. Findings indicate that discrepancies between schemas stem primarily from differing modeling philosophies rather than expressive capacity. While optimized cross-schema mappings significantly improve metadata quality, the diversity of repository workflows also critically affects record completeness. Building on these insights, the study proposes a strategy that leverages the complementary strengths of domain-specific and generic schemas, offering practical guidance for fostering interdisciplinary data sharing.

DataCitedisciplinary metadatametadata interoperability

Latest Papers

What's happening recently
View more

Existing Model Cards and Data Cards describe only static models and datasets, lacking documentation of the execution context surrounding generation, transformation, and evaluation processes—thereby limiting reproducibility and bias analysis. This work proposes Workflow Cards, which extend the structured documentation paradigm to dynamic workflow executions for the first time. Built upon provenance data, Workflow Cards generate machine-readable, structured summaries interpretable by both humans and large language models (LLMs), and incorporate a template designed to answer typical execution-related questions. Experimental results demonstrate that Workflow Cards substantially enhance understanding of workflow executions compared to schema-based query interfaces, nearly doubling answer quality and achieving superior performance under both LLM-as-a-Judge and human evaluations.

Data CardsModel Cardsprovenance data

Community-driven scientific workflow ecosystems often struggle to sustain themselves due to ambiguous maintenance and user support mechanisms, particularly in cross-platform collaboration and heterogeneous execution environments. This study presents the first cross-platform empirical analysis of the nf-core ecosystem, systematically examining 15,760 GitHub issues, 35,411 pull requests, and 895 forum discussions. By integrating metadata and textual features into predictive models, the research uncovers significant disparities in maintenance and support activities across platforms and highlights weak explicit linkages among them. The findings reveal that issues, pull requests, and forum posts predominantly serve distinct roles—coordinating maintenance, facilitating code integration, and providing user support, respectively. Moreover, issue actionability, diagnostic evidence, and depth of interaction emerge as critical determinants of resolution efficiency.

community-drivenheterogeneous execution environmentsmaintenance

Current data processing pipelines for post-training large language models—encompassing cleaning, deduplication, synthesis, and quality filtering—are fragmented and lack auditability and sample-level decision transparency. This work proposes the first end-to-end configurable data processing framework that unifies data ingestion, cleaning, LLM-driven synthesis across eight task types, three-tiered quality gating, and export modules. The system introduces sample-level provenance tracking and a precise hallucination verification mechanism. It supports six input formats and over 100 model APIs via LiteLLM, offering both a YAML-driven command-line interface and a Python API. Outputs are compatible with five training formats used by TRL, Unsloth, and AlignTune, substantially enhancing transparency, reproducibility, and scalability in post-training data preparation.

data curationLLM post-trainingpipeline auditability

Hot Scholars

GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
XL

Xiangtai Li

Research Scientist, Tiktok, SG; MMLab@NTU
Generative AIComputer Vision
CH

Conghui He

Shanghai AI Laboratory
Data-centric AILLMDocument Intelligence
LS

Lifeng Shang

Huawei Noah's Ark Lab
Machine LearningComputer VisionPattern ReconitionNatural Language Processing