Knowledge engineering for open science: Building and deploying knowledge bases for metadata standards

📅 2025-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address insufficient standardization, poor domain adaptability, and practical challenges in implementing FAIR principles for research data metadata in open science, this paper proposes a knowledge engineering–based metadata templating approach. It formalizes domain-specific metadata standards as reusable, logically inferable knowledge templates; designs lightweight, dual-mode (Web form and spreadsheet) acquisition interfaces with real-time validation; and enables cross-platform intelligent integration via declarative knowledge representation. The method has been adopted by multiple international scientific consortia to establish standardized metadata frameworks and has driven the development of over ten data annotation systems. Empirical outcomes demonstrate significant improvements in metadata syntactic and semantic consistency, domain alignment, and cross-platform interoperability. By embedding domain knowledge into machine-processable templates, the approach delivers a scalable, maintainable, knowledge-driven paradigm for FAIR data infrastructure.

Technology Category

Application Category

📝 Abstract
Scientists strive to make their datasets available in open repositories, with the goal that they be findable, accessible, interoperable, and reusable (FAIR). Although it is hard for most investigators to remember all the guiding principles associated with FAIR data, there is one overarching requirement: The data need to be annotated with rich, discipline-specific, standardized metadata. The Center for Expanded Data Annotation and Retrieval (CEDAR) builds technology that enables scientists to encode metadata standards as templates that enumerate the attributes of different kinds of experiments. These metadata templates capture preferences regarding how data should be described and what a third party needs to know to make sense of the datasets. CEDAR templates describing community metadata preferences have been used to standardize metadata for a variety of scientific consortia. They have been used as the basis for data-annotation systems that acquire metadata through Web forms or through spreadsheets, and they can help correct metadata to ensure adherence to standards. Like the declarative knowledge bases that underpinned intelligent systems decades ago, CEDAR templates capture the knowledge in symbolic form, and they allow that knowledge to be applied in a variety of settings. They provide a mechanism for scientific communities to create shared metadata standards and to encode their preferences for the application of those standards, and for deploying those standards in a range of intelligent systems to promote open science.
Problem

Research questions and friction points this paper is trying to address.

Building knowledge bases for FAIR metadata standards
Creating templates to standardize scientific metadata
Deploying metadata standards in intelligent systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

CEDAR templates standardize metadata for experiments
Web forms and spreadsheets acquire metadata efficiently
Symbolic knowledge bases promote FAIR data adherence
M
Mark A. Musen
Stanford Center for Biomedical Informatics Research, Stanford University School of Medicine
M
Martin J. O'Connor
Stanford Center for Biomedical Informatics Research, Stanford University School of Medicine
Josef Hardi
Josef Hardi
Stanford University
Computer Science
M
Marcos Martinez-Romero
Stanford Center for Biomedical Informatics Research, Stanford University School of Medicine