Score
Specifying and packaging licenses, metadata, and provenance so datasets and artifacts can be released transparently and reused legally. This covers documenting modeling assumptions, data-generation procedures, and benchmark artifacts to ensure auditability, reproducibility, and compliance with open-source/open-data terms.
In the context of open science, there is an urgent need to implement the FAIR principles (Findable, Accessible, Interoperable, Reusable) in scientific data management, yet systematic guidance on achieving FAIR compliance through big data software reference architectures (SRAs) remains lacking. Method: We conducted a rigorous systematic literature review, screening 323 publications—including those from authoritative databases and expert recommendations—and performed structured data extraction and evaluation aligned with predefined research questions. Contribution/Results: The study identifies seven generic FAIR-compliant SRAs, thirteen scenario-specific FAIR pipelines, and three fully FAIR-compatible SRAs. It uncovers critical bottlenecks in metadata standardization, cross-platform interoperability, and long-term reusability. Furthermore, it establishes the first classification framework and empirical evaluation system for FAIR-oriented big data SRAs, thereby filling a significant research gap and providing a methodological foundation and strategic direction for future SRA design, policy formulation, and tool development.
This study addresses the lack of transparency and traceability in the training data lifecycle of large language models (LLMs), which undermines their trustworthiness and auditability. Through a systematic review of 95 relevant publications over the past decade, this work proposes the first unified classification framework tailored to the LLM data lifecycle, structured around three core dimensions: data provenance, transparency, and traceability. The framework integrates key technical approaches—including data generation, watermarking, bias measurement, privacy preservation, and governance tools—thereby clarifying the field’s boundaries and revealing inherent trade-offs between transparency and opacity. Furthermore, it synthesizes emerging research trends and open challenges, offering both a theoretical foundation and practical guidance for enhancing the credibility of LLM training data.
Current dataset licensing risk assessment relies heavily on static license terms, failing to address rights erosion and license modifications arising from redistribution; manual evaluation is inherently unscalable. This paper introduces “data-lifecycle-aware compliance” as a novel paradigm and proposes NEXUS, an AI-driven compliance system that enables automated, end-to-end risk identification across the full dataset lifecycle—including redistribution pathways and rights evolution. NEXUS integrates multi-source metadata graph construction, semantic license parsing, collaborative AI agent tracking, and large-scale legal relationship reasoning. Empirical evaluation across 17,429 entities and 8,072 license clauses reveals that only 21% of commercially labeled datasets are actually legally usable; NEXUS achieves significantly higher accuracy and efficiency in compliance judgment than domain-expert human evaluators.
The widespread adoption of open-source machine learning (ML) datasets and models has intensified risks including data poisoning, supply-chain attacks, and regulatory non-compliance. Method: This paper proposes the first verifiable, end-to-end ML provenance framework integrating Trusted Execution Environments (TEEs) and transparent logging—built upon SPDX/SLSA standards and leveraging Intel SGX, hash-chain-based immutable logging, and zero-knowledge proofs. Contribution/Results: The framework enables provable artifact authenticity, auditable end-to-end lineage, and co-guaranteed confidentiality and integrity—without compromising intellectual property rights over data or models. Evaluated on two real-world ML pipelines, it achieves 100% metadata tampering detection, full verifiable traceability from training to deployment, and negligible runtime overhead—demonstrating practical viability for secure, compliant ML operations.
This study addresses persistent challenges in open data publishing within industry–academia–government collaboration, including inefficient data management, barriers to data reuse, weak licensing awareness, and insufficient integration of real and synthetic data. Drawing on in-depth analysis of 13 European collaborative project datasets, statistical examination of metadata from 281,000 datasets on Zenodo, and complementary surveys and inductive reasoning, the study reveals three key empirical findings: (1) data collection planning plays a critical, previously underrecognized role; (2) script documentation is extremely rare (only 2.4% of datasets); and (3) licensing practices are widespread but largely noncompliant. It further provides robust evidence that hybrid real-synthetic or simulation-based datasets hold substantial scientific value. Based on these insights, the study proposes an actionable data management framework and concrete standardization recommendations—aimed at enhancing cross-sectoral data reusability, regulatory compliance, and the maturity of open science practices.
This study addresses the dual challenges confronting Open Government Data (OGD) licensing: ambiguous legal status and inadequate cross-jurisdictional adaptability, revealing its fundamentally policy-driven nature—distinct from commercial licensing or public sharing paradigms. Through the first comparative analysis of OGD license terms across 32 countries, coupled with policy document interpretation and intellectual property law theoretical modeling, we identify critical jurisdictional divergences, including database rights regimes and waivability of moral rights. We innovatively propose a “Policy–Jurisdiction” two-dimensional adaptation framework and derive licensing design principles that jointly ensure legal validity, cross-jurisdictional consistency, and reusability efficacy. The findings elevate OGD licenses from technical appendices to core instruments of information policy, substantially enhancing legal certainty and economic conversion rates of government data assets.
Frequent data copyright infringement during large-scale ML model training, coupled with fragmented assumptions, narrow evaluation scopes, and poor cross-method comparability among existing copyright auditing tools, hinders practical deployment. Method: This paper systematically categorizes intrusive (watermark injection) and non-intrusive (fingerprinting-based) auditing paradigms, and—firstly—establishes a unified analytical framework spanning the entire ML pipeline: data collection, preprocessing, training, and inference. Leveraging full-stack ML modeling and controlled cross-method experiments, it characterizes structural trade-offs across assumptions, stage coverage, and real-world robustness. Contribution/Results: It introduces a deployment-oriented evaluation perspective, synthesizes common limitations, and identifies open challenges. The work delivers a taxonomy reference table and a practical implementation guide, providing both theoretical foundations and actionable technical pathways for developing compliant, deployable, and robust data copyright auditing tools.
This study addresses the widespread issue of “permissive license laundering” in open-source AI ecosystems, where models, datasets, and applications labeled as compliant with permissive licenses such as MIT or Apache-2.0 often lack required license texts, copyright notices, or upstream attributions, thereby introducing legal compliance risks. For the first time, this work quantifies the problem across the entire AI supply chain by combining automated crawling, metadata analysis, and manual verification to audit 3,338 datasets, 6,664 models, and 28,516 applications on Hugging Face and GitHub. The findings reveal that 96.5% of datasets and 95.8% of models are non-compliant, with only 5.75% of downstream applications preserving complete license statements. The study advocates determining license validity through legal documents rather than metadata alone and introduces a reproducible, large-scale compliance auditing framework.
In the era of large language models, dataset copyright protection faces significant challenges, including legal frameworks lagging behind technological development and opaque technical environments. This paper systematically proposes three technical approaches to dataset copyright protection: non-intrusive (e.g., digital watermarking), semi-intrusive (e.g., reversible data perturbation), and highly intrusive (e.g., reversible adversarial examples). It introduces the first comprehensive classification framework covering ownership authentication, usage monitoring, and misuse traceability. By integrating watermark embedding, adversarial sample generation, and model provenance analysis, the approach enables autonomous, law-procedure-free dataset copyright verification. The study rigorously characterizes the applicability boundaries and inherent limitations of each method. Key future research directions are identified as scalability, cross-platform standardization, and ethical compliance.
This study addresses the prevalent issue in computer vision datasets wherein image provenance—such as acquisition parameters and preprocessing steps—is often missing or stored separately, leading to compromised traceability, regulatory compliance, and data reusability. To overcome this limitation, the work proposes a novel approach that leverages JSON-LD (JavaScript Object Notation for Linked Data) to define a structured provenance schema and embeds it directly within image files. This integration ensures an inseparable binding between metadata and the image payload, preserving provenance integrity and persistence across workflows. The proposed method maintains compatibility with existing standards while significantly enhancing data maintainability, system interoperability, and the trustworthiness of downstream models trained on such annotated datasets.
This study presents the first systematic assessment of registration and sharing practices for publicly funded research software in the UK, revealing critical challenges: low traceability, poor sharing rates, and severe link rot (45% of URLs invalid or missing). Methodologically, we analyzed metadata from the UK Research and Innovation (UKRI) Gateway to Research (GtR) database, complemented by automated URL validation, platform-specific classification, and hosting-source attribution. Results indicate that software outputs are markedly underreported relative to other research outputs; only 18% are hosted on GitHub, and 25% lack any accessible links. The findings expose a significant policy implementation gap and an absence of dedicated stewardship for research software as scientific infrastructure. This work provides empirical evidence and a foundational benchmark to inform the development of robust research software management frameworks—enhancing long-term reusability, sustainability, and preservation of computational research assets.
This study addresses critical challenges in social impact data sharing—including the absence of standardized metadata, weak interoperability, and low engagement by official statistical agencies—by proposing and implementing an intelligent metadata framework for Social Impact Data Commons. Grounded in FAIR principles, the framework integrates structured metadata standards (e.g., DCAT, Schema.org) with semantic interoperability techniques to enable automated metadata generation, dynamic quality assessment, and cross-domain collaborative management. Empirical validation across multiple use cases demonstrates significant improvements in data discoverability, accessibility, interoperability, and reusability. Its core contribution is the first incorporation of “executable metadata” into a data commons architecture, establishing an evaluable and extensible standardization pathway. This enhances official statistical agencies’ understanding of and participation in innovative data initiatives, thereby providing FAIR-compliant infrastructure for social impact research.