open dataset curation

Designs and builds openly licensed, versioned, and documented data collections for public use by assembling, cleaning, labeling, and organizing source records into reusable datasets. Produces release artifacts, metadata, provenance and licensing information and maintains distribution and governance (versioning, updates, and contribution processes) so others can access and reuse the data reliably.

opendatasetcuration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

The hunt for research data: Development of an open-source workflow for tracking institutionally-affiliated research data publications

Jul 01, 2025
BM
Bryan M. Gee
🏛️ University of Texas Libraries | The University of Texas at Austin

Institutions face significant challenges in systematically tracking their affiliated research data publications, primarily due to the widespread absence, inconsistency, or non-standardization of institutional attribution metadata (e.g., missing or ambiguous institutional names, lack of persistent identifiers such as DOIs) in existing data repositories. Method: We propose the first open-source, institution-centric workflow for tracking data publications, integrating over 70 open APIs and employing multi-source metadata harvesting, normalization, and automated aggregation—thereby reducing reliance on explicit attribution signals like DOIs or manually curated affiliations. Contribution/Results: The workflow enables efficient discovery and consolidation of over 4,000 cross-platform datasets. Evaluation demonstrates substantial improvements in institutional data discoverability, coverage breadth, and retrieval efficiency. This work delivers a reusable technical infrastructure to support research administration and data governance at institutional and systemic levels.

Address challenges in discovering institution-affiliated datasetsDevelop open-source workflow for tracking institutional research dataImprove metadata standardization for comprehensive data retrieval

Insights from Publishing Open Data in Industry-Academia Collaboration

Jan 24, 2025
PS
P. Strandberg
🏛️ Westermo Network Technologies AB | Johannes Kepler University Linz | Silicon Austria Labs GmbH | VTT Technical Research Centre of Finland Ltd.

This study addresses persistent challenges in open data publishing within industry–academia–government collaboration, including inefficient data management, barriers to data reuse, weak licensing awareness, and insufficient integration of real and synthetic data. Drawing on in-depth analysis of 13 European collaborative project datasets, statistical examination of metadata from 281,000 datasets on Zenodo, and complementary surveys and inductive reasoning, the study reveals three key empirical findings: (1) data collection planning plays a critical, previously underrecognized role; (2) script documentation is extremely rare (only 2.4% of datasets); and (3) licensing practices are widespread but largely noncompliant. It further provides robust evidence that hybrid real-synthetic or simulation-based datasets hold substantial scientific value. Based on these insights, the study proposes an actionable data management framework and concrete standardization recommendations—aimed at enhancing cross-sectoral data reusability, regulatory compliance, and the maturity of open science practices.

Data ManagementData SharingData Utilization

How are research data referenced? The use case of the research data repository RADAR

May 13, 2025
DS
Dorothea Strecker
🏛️ Humboldt-Universität zu Berlin | FIZ Karlsruhe

Assessing data reuse and citation practices remains challenging due to fragmented metrics and inconsistent reporting. Method: This study systematically analyzes citation patterns of datasets published in the German RADAR repository, integrating and cross-validating evidence from Google Scholar, DataCite Event Data, and the Data Citation Corpus. It distinguishes “formal citations” (e.g., in reference lists or data availability statements) from “substantive reuse” (i.e., independent scholarly use without author overlap), identified via institutional affiliation matching. Contribution/Results: Among all RADAR datasets, 27.9% received at least one formal citation, of which 21.4% adhered to community citation standards; only 21 datasets (0.5%) evidenced unambiguous external reuse—indicating nascent data reuse maturity. The study establishes a methodological framework for empirically evaluating data impact and advancing robust, interoperable data citation ecosystems.

Assessing data citation practices and their impact metricsExamining dataset referencing frequency and forms in RADARInvestigating data reuse through authorship and reference patterns

Research Data in Scientific Publications: A Cross-Field Analysis

Feb 03, 2025
PY
Puyu Yang
🏛️ University of Amsterdam | University of Copenhagen | University of Bologna

This study identifies a structural imbalance in interdisciplinary data sharing: high reuse rates in STEM fields contrast sharply with low adoption in humanities and social sciences, while persistent undercitation of datasets impedes evidence-based policy and infrastructure development. Leveraging full-text PubMed articles, we construct the first multidisciplinary dataset—enabling simultaneous identification of data mentions and classification of data-related intents—by integrating natural language processing, full-text pattern recognition, cross-disciplinary bibliometrics, and time-series modeling. Key findings include: (1) a marked acceleration in data publication post-2012; (2) highest data publishing activity in business/management and creative arts, yet highest reuse in biological and agricultural sciences; and (3) consistently low dataset citation rates, revealing critical bottlenecks in discoverability and format interoperability. These empirically grounded insights advance data governance frameworks and support formal recognition of datasets as independent scholarly outputs.

Data InfrastructureData SharingInterdisciplinary Differences

Licensing Open Government Data

May 08, 2017
JL
Jyh-An Lee
🏛️ The Chinese University of Hong Kong | Stanford University | Harvard University

This study addresses the dual challenges confronting Open Government Data (OGD) licensing: ambiguous legal status and inadequate cross-jurisdictional adaptability, revealing its fundamentally policy-driven nature—distinct from commercial licensing or public sharing paradigms. Through the first comparative analysis of OGD license terms across 32 countries, coupled with policy document interpretation and intellectual property law theoretical modeling, we identify critical jurisdictional divergences, including database rights regimes and waivability of moral rights. We innovatively propose a “Policy–Jurisdiction” two-dimensional adaptation framework and derive licensing design principles that jointly ensure legal validity, cross-jurisdictional consistency, and reusability efficacy. The findings elevate OGD licenses from technical appendices to core instruments of information policy, substantially enhancing legal certainty and economic conversion rates of government data assets.

Adapting licenses to IP regimesAmbiguous legal status of open dataLegal issues in open government data licenses

Latest Papers

What's happening recently
View more

This study addresses the widespread practice of directly copying open-source code to bypass dependency management, which obscures license compliance risks. Leveraging the World of Code dataset, the authors construct a code reuse network through large-scale clone detection and quantify, for the first time at the scale of the entire open-source ecosystem, the compliance risks arising from such copy-paste reuse. Their analysis reveals that 39.4% of project compositions entail potential license conflicts, yet conventional dependency analysis tools capture only 2.43% of these instances, indicating severe under-detection. Integrating network modeling and regression analysis, the study further finds that code under permissive licenses such as MIT and Apache is reused across programming languages more frequently, whereas public-domain-licensed code exhibits comparatively lower reuse rates.

code reusecopy-based reusecopyright

Hot Scholars

MF

Marzieh Fadaee

Staff Research Scientist, Cohere Labs
Computational LinguisticsMachine LearningNatural Language ProcessingMultilingual NLP
SL

Shayne Longpre

MIT, Stanford, Apple
Deep LearningNatural Language Understanding
JG

Jiahui Geng

Mohamed bin Zayed University of Artificial Intelligence
Artificial IntelligenceNatural Language Processing
JJ

Jinyuan Jia

Assistant Professor, Penn State
AI Security