Score
Designs and implements processes, tools, and documentation to select, verify, and enforce appropriate licenses for datasets and related artifacts; this includes packaging datasets for release, recording provenance and license metadata, detecting and filtering incompatible or unapproved licenses, and applying license-compatibility rules to ensure compliance across artifacts.
Current dataset licensing risk assessment relies heavily on static license terms, failing to address rights erosion and license modifications arising from redistribution; manual evaluation is inherently unscalable. This paper introduces “data-lifecycle-aware compliance” as a novel paradigm and proposes NEXUS, an AI-driven compliance system that enables automated, end-to-end risk identification across the full dataset lifecycle—including redistribution pathways and rights evolution. NEXUS integrates multi-source metadata graph construction, semantic license parsing, collaborative AI agent tracking, and large-scale legal relationship reasoning. Empirical evaluation across 17,429 entities and 8,072 license clauses reveals that only 21% of commercially labeled datasets are actually legally usable; NEXUS achieves significantly higher accuracy and efficiency in compliance judgment than domain-expert human evaluators.
This study addresses the widespread issue of “permissive license laundering” in open-source AI ecosystems, where models, datasets, and applications labeled as compliant with permissive licenses such as MIT or Apache-2.0 often lack required license texts, copyright notices, or upstream attributions, thereby introducing legal compliance risks. For the first time, this work quantifies the problem across the entire AI supply chain by combining automated crawling, metadata analysis, and manual verification to audit 3,338 datasets, 6,664 models, and 28,516 applications on Hugging Face and GitHub. The findings reveal that 96.5% of datasets and 95.8% of models are non-compliant, with only 5.75% of downstream applications preserving complete license statements. The study advocates determining license validity through legal documents rather than metadata alone and introduces a reproducible, large-scale compliance auditing framework.
License identification in open-source software supply chains faces challenges of scale, heterogeneous reuse, and dynamic evolution. To address this, we introduce the first large-scale, temporally annotated, fine-grained license identification dataset. Leveraging the World of Code infrastructure, we scan files containing “license” in their paths; then apply the Winnowing algorithm combined with SPDX standards for approximate matching, identifying 5.5 million distinct license snippets. We further construct a project-to-license (P2L) temporal mapping covering the entire GitHub commit history. Our proposed scalable identification paradigm integrates path-based heuristics with text-based approximate matching. Evaluated via stratified sampling and manual validation, it achieves 92.08% accuracy and an F1-score of 91.11%. The dataset is publicly released to support compliance auditing, license evolution analysis, and tool development.
Frequent data copyright infringement during large-scale ML model training, coupled with fragmented assumptions, narrow evaluation scopes, and poor cross-method comparability among existing copyright auditing tools, hinders practical deployment. Method: This paper systematically categorizes intrusive (watermark injection) and non-intrusive (fingerprinting-based) auditing paradigms, and—firstly—establishes a unified analytical framework spanning the entire ML pipeline: data collection, preprocessing, training, and inference. Leveraging full-stack ML modeling and controlled cross-method experiments, it characterizes structural trade-offs across assumptions, stage coverage, and real-world robustness. Contribution/Results: It introduces a deployment-oriented evaluation perspective, synthesizes common limitations, and identifies open challenges. The work delivers a taxonomy reference table and a practical implementation guide, providing both theoretical foundations and actionable technical pathways for developing compliant, deployable, and robust data copyright auditing tools.
Scientists frequently record experimental metadata in spreadsheets, yet ensuring consistency and standards compliance remains challenging. This paper introduces a spreadsheet-native metadata governance paradigm: customized Excel/CSV templates embed HuBMAP standards; OWL/SKOS ontology-driven controlled vocabularies are integrated; and a web-based real-time semantic validation tool enables immediate, on-entry verification. The approach seamlessly incorporates semantic constraints into familiar spreadsheet workflows—requiring no platform switching or new system adoption. Deployed across the HuBMAP Consortium, it significantly improved multi-omics metadata compliance rates, increased data entry efficiency, and reduced error identification and correction time by over 70%. To our knowledge, this is the first work to deeply embed ontology-based constraints and real-time semantic validation directly within spreadsheet environments, establishing a scalable, practical paradigm for biomedical metadata standardization.
This study presents the first large-scale empirical investigation into the erosion of licensing obligations across AI supply chains, focusing on the phenomenon of license “laundering”—where licenses are either omitted or altered—during the flow from datasets to models and downstream applications. By tracing 232,270 end-to-end supply chain paths on Hugging Face and GitHub, the authors find that 62.3% of chains contain at least one component lacking any declared license. End-to-end retention rates for copyleft licenses fall below 7%, in stark contrast to 95.1% for permissive licenses. The findings expose significant compliance risks and offer actionable governance recommendations for developers, platform operators, and rights holders to strengthen license adherence throughout the AI development lifecycle.
This work addresses the challenges posed by the heterogeneous multimodal nature of enterprise policy documents, which often cause large language models to hallucinate, disrupt table structures, and lack end-to-end controllability—resulting in labor-intensive manual processing requiring 2–3 days per document. To overcome these limitations, the authors propose a governed multi-agent collaboration framework grounded in a shared, versioned rule repository. The framework integrates large language models (LLMs), vision-language models (VLMs), schema validation, and human-in-the-loop mechanisms through six specialized agents that collaboratively perform parsing, multimodal extraction, consistency verification, evaluation, iterative refinement, and personalized artifact generation, while ensuring full traceability across the pipeline. Evaluated on 120 real-world documents, the approach achieves a 96% success rate, automatically extracts 3,896 rules (71.4% auto-approved), produces 812 deployable artifacts, and reduces per-document processing time to 40–125 minutes.