data licensing

Specifying and packaging licenses, metadata, and provenance so datasets and artifacts can be released transparently and reused legally. This covers documenting modeling assumptions, data-generation procedures, and benchmark artifacts to ensure auditability, reproducibility, and compliance with open-source/open-data terms.

datalicensing

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

This study addresses the lack of transparency and traceability in the training data lifecycle of large language models (LLMs), which undermines their trustworthiness and auditability. Through a systematic review of 95 relevant publications over the past decade, this work proposes the first unified classification framework tailored to the LLM data lifecycle, structured around three core dimensions: data provenance, transparency, and traceability. The framework integrates key technical approaches—including data generation, watermarking, bias measurement, privacy preservation, and governance tools—thereby clarifying the field’s boundaries and revealing inherent trade-offs between transparency and opacity. Furthermore, it synthesizes emerging research trends and open challenges, offering both a theoretical foundation and practical guidance for enhancing the credibility of LLM training data.

data provenancelarge language modelstraceability

Must-Read Papers

Most classic and influential ideas
View more

Current dataset licensing risk assessment relies heavily on static license terms, failing to address rights erosion and license modifications arising from redistribution; manual evaluation is inherently unscalable. This paper introduces “data-lifecycle-aware compliance” as a novel paradigm and proposes NEXUS, an AI-driven compliance system that enables automated, end-to-end risk identification across the full dataset lifecycle—including redistribution pathways and rights evolution. NEXUS integrates multi-source metadata graph construction, semantic license parsing, collaborative AI agent tracking, and large-scale legal relationship reasoning. Empirical evaluation across 17,429 entities and 8,072 license clauses reveals that only 21% of commercially labeled datasets are actually legally usable; NEXUS achieves significantly higher accuracy and efficiency in compliance judgment than domain-expert human evaluators.

AI-powered system ensures accurate dataset complianceAssessing dataset legal risk requires lifecycle tracingManual legal compliance is inefficient at scale

Atlas: A Framework for ML Lifecycle Provenance&Transparency

Feb 26, 2025
MS
Marcin Spoczynski
🏛️ Intel Labs

The widespread adoption of open-source machine learning (ML) datasets and models has intensified risks including data poisoning, supply-chain attacks, and regulatory non-compliance. Method: This paper proposes the first verifiable, end-to-end ML provenance framework integrating Trusted Execution Environments (TEEs) and transparent logging—built upon SPDX/SLSA standards and leveraging Intel SGX, hash-chain-based immutable logging, and zero-knowledge proofs. Contribution/Results: The framework enables provable artifact authenticity, auditable end-to-end lineage, and co-guaranteed confidentiality and integrity—without compromising intellectual property rights over data or models. Evaluated on two real-world ML pipelines, it achieves 100% metadata tampering detection, full verifiable traceability from training to deployment, and negligible runtime overhead—demonstrating practical viability for secure, compliant ML operations.

Addresses risks in ML lifecycle transparencyBalances regulatory needs with confidentialityEnhances metadata integrity and data security

Insights from Publishing Open Data in Industry-Academia Collaboration

Jan 24, 2025
PS
P. Strandberg
🏛️ Westermo Network Technologies AB | Johannes Kepler University Linz | Silicon Austria Labs GmbH | VTT Technical Research Centre of Finland Ltd.

This study addresses persistent challenges in open data publishing within industry–academia–government collaboration, including inefficient data management, barriers to data reuse, weak licensing awareness, and insufficient integration of real and synthetic data. Drawing on in-depth analysis of 13 European collaborative project datasets, statistical examination of metadata from 281,000 datasets on Zenodo, and complementary surveys and inductive reasoning, the study reveals three key empirical findings: (1) data collection planning plays a critical, previously underrecognized role; (2) script documentation is extremely rare (only 2.4% of datasets); and (3) licensing practices are widespread but largely noncompliant. It further provides robust evidence that hybrid real-synthetic or simulation-based datasets hold substantial scientific value. Based on these insights, the study proposes an actionable data management framework and concrete standardization recommendations—aimed at enhancing cross-sectoral data reusability, regulatory compliance, and the maturity of open science practices.

Data ManagementData SharingData Utilization

Licensing Open Government Data

May 08, 2017
JL
Jyh-An Lee
🏛️ The Chinese University of Hong Kong | Stanford University | Harvard University

This study addresses the dual challenges confronting Open Government Data (OGD) licensing: ambiguous legal status and inadequate cross-jurisdictional adaptability, revealing its fundamentally policy-driven nature—distinct from commercial licensing or public sharing paradigms. Through the first comparative analysis of OGD license terms across 32 countries, coupled with policy document interpretation and intellectual property law theoretical modeling, we identify critical jurisdictional divergences, including database rights regimes and waivability of moral rights. We innovatively propose a “Policy–Jurisdiction” two-dimensional adaptation framework and derive licensing design principles that jointly ensure legal validity, cross-jurisdictional consistency, and reusability efficacy. The findings elevate OGD licenses from technical appendices to core instruments of information policy, substantially enhancing legal certainty and economic conversion rates of government data assets.

Adapting licenses to IP regimesAmbiguous legal status of open dataLegal issues in open government data licenses

SoK: Dataset Copyright Auditing in Machine Learning Systems

Oct 22, 2024
LD
L. Du
🏛️ Xi’an Jiaotong University | Zhejiang University | Vrije Universiteit Amsterdam | Hangzhou Dianzi University

Frequent data copyright infringement during large-scale ML model training, coupled with fragmented assumptions, narrow evaluation scopes, and poor cross-method comparability among existing copyright auditing tools, hinders practical deployment. Method: This paper systematically categorizes intrusive (watermark injection) and non-intrusive (fingerprinting-based) auditing paradigms, and—firstly—establishes a unified analytical framework spanning the entire ML pipeline: data collection, preprocessing, training, and inference. Leveraging full-stack ML modeling and controlled cross-method experiments, it characterizes structural trade-offs across assumptions, stage coverage, and real-world robustness. Contribution/Results: It introduces a deployment-oriented evaluation perspective, synthesizes common limitations, and identifies open challenges. The work delivers a taxonomy reference table and a practical implementation guide, providing both theoretical foundations and actionable technical pathways for developing compliant, deployable, and robust data copyright auditing tools.

Auditing copyright in ML training datasets to prevent unauthorized data useComparing strengths and weaknesses of existing dataset copyright auditing solutionsEvaluating robustness of auditing tools in real-world ML applications

Latest Papers

What's happening recently
View more

This study addresses the widespread issue of “permissive license laundering” in open-source AI ecosystems, where models, datasets, and applications labeled as compliant with permissive licenses such as MIT or Apache-2.0 often lack required license texts, copyright notices, or upstream attributions, thereby introducing legal compliance risks. For the first time, this work quantifies the problem across the entire AI supply chain by combining automated crawling, metadata analysis, and manual verification to audit 3,338 datasets, 6,664 models, and 28,516 applications on Hugging Face and GitHub. The findings reveal that 96.5% of datasets and 95.8% of models are non-compliant, with only 5.75% of downstream applications preserving complete license statements. The study advocates determining license validity through legal documents rather than metadata alone and introduces a reproducible, large-scale compliance auditing framework.

AI supply chainattributionlicense compliance

Dataset Ownership in the Era of Large Language Models

Sep 07, 2025
KL
Kun Li
🏛️ Shandong University

In the era of large language models, dataset copyright protection faces significant challenges, including legal frameworks lagging behind technological development and opaque technical environments. This paper systematically proposes three technical approaches to dataset copyright protection: non-intrusive (e.g., digital watermarking), semi-intrusive (e.g., reversible data perturbation), and highly intrusive (e.g., reversible adversarial examples). It introduces the first comprehensive classification framework covering ownership authentication, usage monitoring, and misuse traceability. By integrating watermark embedding, adversarial sample generation, and model provenance analysis, the approach enables autonomous, law-procedure-free dataset copyright verification. The study rigorously characterizes the applicability boundaries and inherent limitations of each method. Key future research directions are identified as scalability, cross-platform standardization, and ethical compliance.

Addressing unauthorized use in opaque or decentralized environmentsDeveloping technical approaches for dataset ownership verificationEnsuring robust copyright protection for datasets

This study addresses the prevalent issue in computer vision datasets wherein image provenance—such as acquisition parameters and preprocessing steps—is often missing or stored separately, leading to compromised traceability, regulatory compliance, and data reusability. To overcome this limitation, the work proposes a novel approach that leverages JSON-LD (JavaScript Object Notation for Linked Data) to define a structured provenance schema and embeds it directly within image files. This integration ensures an inseparable binding between metadata and the image payload, preserving provenance integrity and persistence across workflows. The proposed method maintains compatibility with existing standards while significantly enhancing data maintainability, system interoperability, and the trustworthiness of downstream models trained on such annotated datasets.

computer vision datasetsdata traceabilityimage provenance

Tracking research software outputs in the UK

Jul 30, 2025
DC
Domhnall Carlin
🏛️ Queen’s University Belfast

This study presents the first systematic assessment of registration and sharing practices for publicly funded research software in the UK, revealing critical challenges: low traceability, poor sharing rates, and severe link rot (45% of URLs invalid or missing). Methodologically, we analyzed metadata from the UK Research and Innovation (UKRI) Gateway to Research (GtR) database, complemented by automated URL validation, platform-specific classification, and hosting-source attribution. Results indicate that software outputs are markedly underreported relative to other research outputs; only 18% are hosted on GitHub, and 25% lack any accessible links. The findings expose a significant policy implementation gap and an absence of dedicated stewardship for research software as scientific infrastructure. This work provides empirical evidence and a foundational benchmark to inform the development of robust research software management frameworks—enhancing long-term reusability, sustainability, and preservation of computational research assets.

Ensuring long-term accessibility of research software for scienceLow artifact sharing and URL errors in research softwareTracking and storing UK research software outputs effectively

Smart Metadata in Action: The Social Impact Data Commons

Nov 21, 2025
JS
Joanna Schroeder
🏛️ University of Virginia | Drexel University

This study addresses critical challenges in social impact data sharing—including the absence of standardized metadata, weak interoperability, and low engagement by official statistical agencies—by proposing and implementing an intelligent metadata framework for Social Impact Data Commons. Grounded in FAIR principles, the framework integrates structured metadata standards (e.g., DCAT, Schema.org) with semantic interoperability techniques to enable automated metadata generation, dynamic quality assessment, and cross-domain collaborative management. Empirical validation across multiple use cases demonstrates significant improvements in data discoverability, accessibility, interoperability, and reusability. Its core contribution is the first incorporation of “executable metadata” into a data commons architecture, establishing an evaluable and extensible standardization pathway. This enhances official statistical agencies’ understanding of and participation in innovative data initiatives, thereby providing FAIR-compliant infrastructure for social impact research.

Developing FAIR data systems using smart metadataEvaluating metadata adherence to FAIR guidelines standardsImplementing actionable metadata for Social Impact Data Commons

Hot Scholars

GK

Gopi Krishnan Rajbahadur

Senior Researcher, Centre for Software Excellence, Huawei
Machine learningSoftware analyticsAI4SESE4AI
AE

Ahmed E. Hassan

Mustafa Prize Laureate, ACM/IEEE/NSERC Steacie Fellow, ACM Influential/IEEE Distinguished Educator
Mining Software RepositoriesSoftware AnalyticsEmpirical Software EngineeringSoftware
AM

Audris Mockus

University of Tennessee
Digital ArchaeologySoftware EngineeringVisualizationOptimization
FJ

Frederik J. Zuiderveen Borgesius

Professor ICT and Law, iHub, Radboud University, The Netherlands
Law and technologyprivacydata protection lawnon-discrimination law
CS

Cristiana Santos

Utrecht University
Compliance with Data Protection LawDark PatternsTracking