Score
Designs and builds openly licensed, versioned, and documented data collections for public use by assembling, cleaning, labeling, and organizing source records into reusable datasets. Produces release artifacts, metadata, provenance and licensing information and maintains distribution and governance (versioning, updates, and contribution processes) so others can access and reuse the data reliably.
Existing dataset documentation tools struggle to achieve real-world adoption due to ambiguous value propositions, misalignment with practical contexts, insufficient attention to human labor costs, and a lack of systemic integration. This study addresses these challenges through a mixed-methods systematic scoping review of 59 relevant publications, combining qualitative coding with quantitative synthesis to uncover the underlying motivations driving tool design and their relationship to institutional norms. The analysis identifies four key patterns that hinder adoption and advances a responsible AI design perspective that shifts emphasis from individual accountability to institutional solutions. The work advocates embedding sustainable documentation practices within organizational workflows and cultures, offering the HCI community actionable pathways toward institutionalizing responsible data stewardship.
Institutions face significant challenges in systematically tracking their affiliated research data publications, primarily due to the widespread absence, inconsistency, or non-standardization of institutional attribution metadata (e.g., missing or ambiguous institutional names, lack of persistent identifiers such as DOIs) in existing data repositories. Method: We propose the first open-source, institution-centric workflow for tracking data publications, integrating over 70 open APIs and employing multi-source metadata harvesting, normalization, and automated aggregation—thereby reducing reliance on explicit attribution signals like DOIs or manually curated affiliations. Contribution/Results: The workflow enables efficient discovery and consolidation of over 4,000 cross-platform datasets. Evaluation demonstrates substantial improvements in institutional data discoverability, coverage breadth, and retrieval efficiency. This work delivers a reusable technical infrastructure to support research administration and data governance at institutional and systemic levels.
This study addresses persistent challenges in open data publishing within industry–academia–government collaboration, including inefficient data management, barriers to data reuse, weak licensing awareness, and insufficient integration of real and synthetic data. Drawing on in-depth analysis of 13 European collaborative project datasets, statistical examination of metadata from 281,000 datasets on Zenodo, and complementary surveys and inductive reasoning, the study reveals three key empirical findings: (1) data collection planning plays a critical, previously underrecognized role; (2) script documentation is extremely rare (only 2.4% of datasets); and (3) licensing practices are widespread but largely noncompliant. It further provides robust evidence that hybrid real-synthetic or simulation-based datasets hold substantial scientific value. Based on these insights, the study proposes an actionable data management framework and concrete standardization recommendations—aimed at enhancing cross-sectoral data reusability, regulatory compliance, and the maturity of open science practices.
Assessing data reuse and citation practices remains challenging due to fragmented metrics and inconsistent reporting. Method: This study systematically analyzes citation patterns of datasets published in the German RADAR repository, integrating and cross-validating evidence from Google Scholar, DataCite Event Data, and the Data Citation Corpus. It distinguishes “formal citations” (e.g., in reference lists or data availability statements) from “substantive reuse” (i.e., independent scholarly use without author overlap), identified via institutional affiliation matching. Contribution/Results: Among all RADAR datasets, 27.9% received at least one formal citation, of which 21.4% adhered to community citation standards; only 21 datasets (0.5%) evidenced unambiguous external reuse—indicating nascent data reuse maturity. The study establishes a methodological framework for empirically evaluating data impact and advancing robust, interoperable data citation ecosystems.
This study identifies a structural imbalance in interdisciplinary data sharing: high reuse rates in STEM fields contrast sharply with low adoption in humanities and social sciences, while persistent undercitation of datasets impedes evidence-based policy and infrastructure development. Leveraging full-text PubMed articles, we construct the first multidisciplinary dataset—enabling simultaneous identification of data mentions and classification of data-related intents—by integrating natural language processing, full-text pattern recognition, cross-disciplinary bibliometrics, and time-series modeling. Key findings include: (1) a marked acceleration in data publication post-2012; (2) highest data publishing activity in business/management and creative arts, yet highest reuse in biological and agricultural sciences; and (3) consistently low dataset citation rates, revealing critical bottlenecks in discoverability and format interoperability. These empirically grounded insights advance data governance frameworks and support formal recognition of datasets as independent scholarly outputs.
This study addresses the dual challenges confronting Open Government Data (OGD) licensing: ambiguous legal status and inadequate cross-jurisdictional adaptability, revealing its fundamentally policy-driven nature—distinct from commercial licensing or public sharing paradigms. Through the first comparative analysis of OGD license terms across 32 countries, coupled with policy document interpretation and intellectual property law theoretical modeling, we identify critical jurisdictional divergences, including database rights regimes and waivability of moral rights. We innovatively propose a “Policy–Jurisdiction” two-dimensional adaptation framework and derive licensing design principles that jointly ensure legal validity, cross-jurisdictional consistency, and reusability efficacy. The findings elevate OGD licenses from technical appendices to core instruments of information policy, substantially enhancing legal certainty and economic conversion rates of government data assets.
This study addresses the widespread practice of directly copying open-source code to bypass dependency management, which obscures license compliance risks. Leveraging the World of Code dataset, the authors construct a code reuse network through large-scale clone detection and quantify, for the first time at the scale of the entire open-source ecosystem, the compliance risks arising from such copy-paste reuse. Their analysis reveals that 39.4% of project compositions entail potential license conflicts, yet conventional dependency analysis tools capture only 2.43% of these instances, indicating severe under-detection. Integrating network modeling and regression analysis, the study further finds that code under permissive licenses such as MIT and Apache is reused across programming languages more frequently, whereas public-domain-licensed code exhibits comparatively lower reuse rates.