Score
Designs and implements machine-readable datasets and metadata using Linked Data principles and open-data practices, including encoding ontology networks and controlled vocabularies to enable semantic interoperability. Builds and publishes data and metadata with clear provenance, licensing, and access mechanisms so others can discover, reuse, and reproduce analyses or test claims.
To address insufficient standardization, poor domain adaptability, and practical challenges in implementing FAIR principles for research data metadata in open science, this paper proposes a knowledge engineering–based metadata templating approach. It formalizes domain-specific metadata standards as reusable, logically inferable knowledge templates; designs lightweight, dual-mode (Web form and spreadsheet) acquisition interfaces with real-time validation; and enables cross-platform intelligent integration via declarative knowledge representation. The method has been adopted by multiple international scientific consortia to establish standardized metadata frameworks and has driven the development of over ten data annotation systems. Empirical outcomes demonstrate significant improvements in metadata syntactic and semantic consistency, domain alignment, and cross-platform interoperability. By embedding domain knowledge into machine-processable templates, the approach delivers a scalable, maintainable, knowledge-driven paradigm for FAIR data infrastructure.
Although scientific data increasingly adhere to the FAIR principles and employ standardized identifiers, practical interoperability remains hindered by heterogeneity in identifier systems and data models. This work proposes and implements two synergistic tools—Babel and ORION—to bridge this gap. Babel constructs clusters of equivalent identifiers through mapping-based clustering and exposes them via a high-performance quantitative API, while ORION standardizes heterogeneous knowledge bases by aligning them to a community-governed common data model. Together, they systematically address the longstanding disconnect between the FAIR “Interoperable” principle and its real-world implementation. The integration of these tools has enabled the construction of a fully interoperable knowledge base, substantially enhancing cross-resource data integration and query capabilities. The resulting framework is publicly available.
This study addresses persistent challenges in open data publishing within industry–academia–government collaboration, including inefficient data management, barriers to data reuse, weak licensing awareness, and insufficient integration of real and synthetic data. Drawing on in-depth analysis of 13 European collaborative project datasets, statistical examination of metadata from 281,000 datasets on Zenodo, and complementary surveys and inductive reasoning, the study reveals three key empirical findings: (1) data collection planning plays a critical, previously underrecognized role; (2) script documentation is extremely rare (only 2.4% of datasets); and (3) licensing practices are widespread but largely noncompliant. It further provides robust evidence that hybrid real-synthetic or simulation-based datasets hold substantial scientific value. Based on these insights, the study proposes an actionable data management framework and concrete standardization recommendations—aimed at enhancing cross-sectoral data reusability, regulatory compliance, and the maturity of open science practices.
Digital humanities face challenges in provenance tracking and change management for cultural heritage metadata, as existing RDF-based approaches suffer from weak standards compliance (e.g., W3C RDF reification, n-ary relations) and poor cross-domain interoperability. Method: This study conducts a systematic, multidimensional empirical evaluation of six mainstream semantic models—Named Graphs, RDF*, PROV-O, among others—assessing their standards conformance, extensibility, and domain adaptability specifically within cultural heritage contexts. Contribution/Results: We propose a practice-oriented provenance modeling selection framework that explicitly characterizes trade-offs among trustworthiness assurance, computational overhead, and interoperability. The framework delivers reusable, verifiable decision support for metadata provenance modeling in digital humanities projects, thereby bridging a critical methodological gap in the deep adaptation of Semantic Web technologies to humanities scholarship.
This study addresses the low accuracy of large language models (LLMs) in FAIR-compliance validation of biosample metadata. We propose a structured-knowledge-guided prompting method, integrating the CEDAR template repository, domain-specific data dictionaries, and GPT-4 to construct a metadata standards-conformance verification framework—demonstrated on human lung cancer biosamples. Experimental results show that incorporating structured knowledge significantly improves field-level standards compliance from 79% to 97% (p < 0.01), providing the first empirical evidence that structured knowledge bases can overcome performance bottlenecks inherent to purely text-based LLM prompting in metadata governance. Our approach establishes a novel paradigm for automated, high-accuracy, and interpretable FAIR metadata quality control, enabling scalable, standards-aware curation of biomedical metadata.
This work addresses the lack of semantic interoperability in structured data (e.g., JSON, YAML, CSV) within scientific workflows, which hinders consistent interpretation. The authors introduce an RDF authoring view in the MetaConfigurator editor that leverages AI-assisted generation of RML mappings to automatically transform structured data into RDF. The system supports triple editing, SPARQL querying, and knowledge graph visualization. Key innovations include the first integration of large language models for natural language-to-SPARQL translation, bidirectional synchronization between JSON-LD and RDF triples, and ontology-aware IRI auto-completion. Demonstrated on MOF synthesis experiments, the approach successfully converts JSON protocols into semantic knowledge graphs, enabling interactive exploration of relationships between experimental conditions and outcomes, thereby significantly lowering the barrier to adopting Semantic Web technologies.
This study investigates how to enhance retrieval effectiveness over RDF datasets while preserving the semantic faithfulness of metadata generated by large language models (LLMs) to the original data. It formulates metadata generation for the first time as a system-level information retrieval problem and systematically evaluates six LLM-based strategies—ranging from unconstrained rewriting to knowledge graph–grounded agent approaches—in terms of the trade-off between retrieval performance and content faithfulness. Experimental results show that unconstrained rewriting yields the greatest retrieval gains but suffers from the lowest faithfulness, whereas profile-guided rewriting achieves the best balance between the two objectives. The work reveals that retrieval improvements may stem from unfaithful semantic expansions and proposes a new paradigm that jointly optimizes effectiveness and trustworthiness in metadata generation for semantic data retrieval.
This work addresses the limitations of existing ontology documentation tools in supporting modular modeling and human readability, particularly in handling cross-module entities and annotations. To overcome these challenges, the authors refactor and extend the LODE framework by introducing a modular Reader-Model-Viewer architecture that decouples parsing, modeling, and rendering components. Implemented as a web service, the new framework provides enhanced capabilities for generating OWL ontology documentation, featuring dedicated entity pages, RDF provenance tracking, and Markdown-based rendering. These improvements significantly increase the intelligibility and reusability of modular scientific knowledge graph ontologies. The framework has been successfully applied to the documentation of the SKG-O ontology, demonstrating its practical utility and effectiveness.
This work proposes HERITRACE, a novel system that integrates domain-expert-driven interactive editing, fully automated change snapshots, and temporal rollback mechanisms into RDF data curation to address semantic ambiguities and errors—such as duplicate records and metadata conflicts. Built upon SHACL validation rules and YAML-based display configurations, HERITRACE automatically generates intuitive form-based interfaces that enable experts to efficiently and traceably edit RDF triples while comprehensively logging provenance and revision history. Evaluated in an OpenCitations Meta bibliographic deduplication scenario, the system successfully assisted experts in identifying DOI conflicts, correcting erroneous merges, and reverting to prior states, thereby demonstrating its effectiveness and practicality in supporting auditable, recoverable semantic data governance.