Score
Designs and implements resources and processes to collect, standardize, map, and maintain terminologies and controlled vocabularies; this includes defining and applying normalization rules for spelling, morphology, casing, synonyms, abbreviations and variant forms, and building tooling for lookup, validation, versioning, and integration with downstream systems.
Scientific data often suffers from poor discoverability, limited sharing, inefficient reuse, and high curation costs due to inadequate standardization of fields and terminology. To address these challenges, this paper proposes a FAIR-aligned standardization framework. Methodologically, it integrates structured vocabulary design, context-aware metadata modeling, and data homogenization strategies to systematically mitigate semantic noise and concept explosion. Crucially, it embeds ten actionable, principle-based rules into a dynamic, evolving data governance process. Empirical evaluation demonstrates that the framework significantly improves metadata quality and semantic consistency, reduces data management overhead, and enhances data findability, interoperability, and long-term reusability—thereby enabling robust, real-world implementation of the FAIR principles in scientific research settings.
This study addresses the low accuracy of large language models (LLMs) in FAIR-compliance validation of biosample metadata. We propose a structured-knowledge-guided prompting method, integrating the CEDAR template repository, domain-specific data dictionaries, and GPT-4 to construct a metadata standards-conformance verification framework—demonstrated on human lung cancer biosamples. Experimental results show that incorporating structured knowledge significantly improves field-level standards compliance from 79% to 97% (p < 0.01), providing the first empirical evidence that structured knowledge bases can overcome performance bottlenecks inherent to purely text-based LLM prompting in metadata governance. Our approach establishes a novel paradigm for automated, high-accuracy, and interpretable FAIR metadata quality control, enabling scalable, standards-aware curation of biomedical metadata.
Although scientific data increasingly adhere to the FAIR principles and employ standardized identifiers, practical interoperability remains hindered by heterogeneity in identifier systems and data models. This work proposes and implements two synergistic tools—Babel and ORION—to bridge this gap. Babel constructs clusters of equivalent identifiers through mapping-based clustering and exposes them via a high-performance quantitative API, while ORION standardizes heterogeneous knowledge bases by aligning them to a community-governed common data model. Together, they systematically address the longstanding disconnect between the FAIR “Interoperable” principle and its real-world implementation. The integration of these tools has enabled the construction of a fully interoperable knowledge base, substantially enhancing cross-resource data integration and query capabilities. The resulting framework is publicly available.
This work addresses the need for automated digital rights policy generation in multi-institutional, culturally oriented trusted data spaces. We propose a large language model (LLM)-based natural language-to-ODRL policy mapping method. Our approach uniquely integrates the W3C ODRL ontology and its structured documentation as core components of prompt engineering to guide GPT-4 in generating high-fidelity, standards-compliant policies. Additionally, we introduce an ontology-adaptation heuristic tailored for knowledge graph construction to enhance semantic alignment. Evaluated on 12 culturally diverse use cases spanning varying complexity levels, our method achieves a policy generation accuracy of 91.95%, significantly outperforming existing baselines. The contribution lies in establishing a scalable, interpretable, and standards-aligned automation paradigm for open digital rights management—bridging natural language requirements with formal, machine-processable ODRL policies.
Heterogeneous metadata across language resources (LRs) impedes their discoverability and interoperability. Method: This study proposes a unified RDF metadata model integrating the DCAT and META-SHARE ontologies—achieving, for the first time, deep semantic coupling between them. It introduces a novel evaluation paradigm grounded in real user queries (CML) and designs an open-vocabulary-driven, API-based mechanism for machine-readable access. Leveraging the Linked Data technology stack and the Linghub portal, the system supports text search, faceted browsing, and SPARQL querying. Contribution/Results: Empirical evaluation demonstrates significant improvements in LR discoverability, accessibility, and subset extractability; effectively identifies critical metadata heterogeneity issues; and validates that standards-driven ontology integration delivers substantial, measurable gains in LR infrastructure interoperability.
This work addresses the limitation of existing text-to-process modeling approaches, which predominantly focus on control flow while neglecting resource and collaboration perspectives, thereby struggling to generate complete multi-party models. To overcome this, the authors propose a resource-aware generative pipeline that systematically incorporates the resource dimension into large language model (LLM)-driven process modeling for the first time. The method automatically constructs BPMN 2.0 collaboration diagrams from natural language descriptions, explicitly capturing organizational pools, role-based lanes, and inter-organizational message events, and employs an orthogonal layout algorithm for automated diagram arrangement. Experimental results across ten business processes and nine LLMs demonstrate that the approach accurately extracts resource-related information, maintains high control-flow quality, and incurs only minimal runtime overhead, advancing generative process modeling toward more collaborative and resource-complete representations.
Biomedical metadata often suffers from incompleteness or noncompliance with community standards, undermining its discoverability, interoperability, and reusability. To address this challenge, this work proposes a large language model (LLM)-based approach for metadata standardization that innovatively integrates ontology-aware constraints with real-time tool invocation. By dynamically querying authoritative terminology services (e.g., OntoPortal) and machine-actionable metadata templates, the LLM retrieves up-to-date specifications on demand rather than relying solely on static training knowledge. Evaluated on 839 legacy HuBMAP records, the method significantly outperforms a baseline LLM-only strategy, achieving higher prediction accuracy and compliance across both controlled and uncontrolled fields.
This study addresses the challenge of structuring complex French regulatory texts on repairability by proposing a two-stage large language model–assisted approach. The method first performs open information extraction fused with embeddings grounded in the SEMLEG core ontology to normalize labels and induce candidate properties; it then leverages these results to guide closed-world triple extraction across the full corpus, constructing an RDF knowledge graph. Innovatively integrating open and closed extraction paradigms, this work is the first in the legal domain to systematically induce and validate predicate signatures with respect to their domain–range constraints. Experimental results demonstrate robust structured output, near-complete class alignment, and a significant reduction in entity and predicate redundancy—introducing new attributes in fewer than 20% of extracted triples—while uncovering novel semantic combinations of existing predicates.
Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.
This work addresses the challenge of reliably translating natural language into industrial-grade, deployable SysMLv2 models. The authors propose an iterative generate-check-repair framework that, for the first time, integrates a production-level SysMLv2 conformance checker directly into the generation process as a control mechanism rather than a post-processing step. By combining large language model (LLM) generation with deterministic diagnostic feedback and targeted repair strategies—and terminating only when zero errors remain—the method achieves perfect compliance. Evaluated across 604 test cases derived from 151 prompts and four distinct LLMs, the approach elevates single-pass generation compliance from 51.16% to 100%, enabling robust, direct translation of natural language specifications into engineering-ready SysMLv2 models.