SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.
📝 Abstract
Scientific processes are often described in heterogeneous article discourse, with details needed for comparison, reproducibility, reuse, and automation dispersed across prose, tables, figures, protocols, and supplementary files. We present the first release of SciSchema.org, a multidisciplinary collection of 16 expert-annotated schemas spanning Biology & Biotechnology, Materials & Chemistry, Imaging & Measurement, Physics, and Psychology. Each schema defines reusable fields for describing process instances, including inputs, outputs, materials, instruments or software, parameters, conditions, procedural steps, measurements, and provenance-related information. The schemas were created through a human-in-the-loop schema-mining workflow in which large language models generated candidate structures from process specifications, scientific articles, and expert feedback, followed by domain-expert construction of final master schemas. The dataset contains final schemas in JSON Schema and SHACL formats, intermediate model-generated schemas, expert-feedback records, source-paper metadata, community-development materials, and analysis scripts. Technical validation assessed schema structure, development provenance, expert review, and syntactic conformance. The collection supports structured annotation, metadata enrichment, scientific knowledge graphs, information extraction, semantic publishing, and cross-study comparison.
Problem

Research questions and friction points this paper is trying to address.

scientific process description
structured representation
reproducibility
heterogeneous data
schema
Innovation

Methods, ideas, or system contributions that make the work stand out.

structured scientific process
schema mining
human-in-the-loop
large language models
scientific knowledge graphs
🔎 Similar Papers
2024-06-08Annual Meeting of the Association for Computational LinguisticsCitations: 2
Jennifer D'Souza
Jennifer D'Souza
TIB Leibniz Information Centre for Science and Technology
Natural Language ProcessingScientific Knowledge ExtractionLLM EvaluationScientometrics
S
Sameer Sadruddin
TIB Leibniz Information Centre for Science and Technology, Hannover, Germany
A
Anisa Rula
University of Brescia, Brescia, Italy
A
Ana Bossler
University of Alicante, Alicante, Spain
A
Andrés Fullana
University of Alicante, Alicante, Spain
E
Enric Bas
University of Alicante, Alicante, Spain
S
Syed Ather
Georgia Institute of Technology, Atlanta, United States
Defne Circi
Defne Circi
Graduate Student, Duke University
NLPMaterials ScienceMLMaterials Informatics
Anlan Chen
Anlan Chen
University of Bath, School of Management
EntrepreneurshipEcosystemsSustainable innovation and growth strategies
L
L. Catherine Brinson
Duke University, Durham, United States
A
Alyssa Columbus
Johns Hopkins University, Baltimore, United States
G
George Demetriou
University of Manchester, Manchester, United Kingdom
D
Dongjun Jeong
Narnia Labs, Daejeon, South Korea
Tarun Kumar
Tarun Kumar
Hewlett Packard Enterprise
Machine LearningComplex Network AnalysisNetwork Science
Frank Krüger
Frank Krüger
Hochschule Wismar
Text MiningResearch Data ManagementData ScienceProvenance
S
Sascha Genehr
Wismar University of Applied Sciences, Wismar, Germany
K
Kai Budde-Sagert
University of Rostock, Rostock, Germany
A
Anamaria Leonescu
University College London, London, United Kingdom
F
Francesco Lodola
University of Milano-Bicocca, Milan, Italy
C
Chiara Florindi
University of Milano-Bicocca, Milan, Italy
G
Gagana Balasubramanya Murthy
Cambridge Institute of Technology, Bengaluru, India
S
Samson Oluwapelumi Olagbile
Dangote Fertiliser Limited, Lagos, Nigeria
N
Nazia Riasat
North Dakota State University, Fargo, United States
Y
Yan Sha
University of Alberta, Edmonton, Canada
K
Kevin Shen
SES AI, Woburn, United States