From Code Archival to Knowledge Graph: Bridging Software Heritage, COAR Notify and Wikidata

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过构建从代码存档到知识图谱的端到端协调管道,利用Wikidata作为连接器,解决了学术记录与源代码之间缺乏验证链接的问题。
📝 Abstract
Software is a first-class scientific object, yet validated links between source code and the scholarly record remain largely absent from the Linked Open Data (LOD) cloud, isolating archived artefacts from semantic discovery. This paper presents an end-to-end reconciliation pipeline that harvests, validates, and models publication-to-repository pairs from sources where the link between a paper and its source code is explicit and editorially verified: the software-centric journals JOSS, SoftwareX, and IPOL, together with the reproducibility reports of the SIGMOD Availability and Reproducibility Initiative (ARI). This yields a curated corpus of 4,397 $\langle$DOI, repository-URL$\rangle$ pairs. We design two distinct application profiles grounded in Wikidata classes (one for scholarly articles, one for software instances) aligned with the schema.org and CodeMeta vocabularies. This architectural separation enables rule-based reconciliation at two granularities: lightweight, inline publication references or standalone, first-class Wikidata software nodes equipped with SWHIDs, Software Heritage's content-addressed identifiers. A read-only lookup against Wikidata shows that only 82 of the harvested repositories were already modelled there; human-reviewed batches have since created 4{,}182 new software items cross-linked to their articles. We further show that payloads of the emerging COAR Notify protocol, an external effort we do not develop, map natively onto our input format, so the same backend could later serve a live enrichment stream. Our core contribution is a pair of application profiles that turn Wikidata into a connector between the scholarly record and archived source code; we openly release all code, application profiles, and harvested datasets.
Problem

Research questions and friction points this paper is trying to address.

Linked Open Data
source code
scholarly record
semantic discovery
Wikidata
Innovation

Methods, ideas, or system contributions that make the work stand out.

end-to-end reconciliation pipeline
Wikidata application profiles
SWHID
COAR Notify protocol
scholarly record and archived source code linkage
🔎 Similar Papers
No similar papers found.
C
Camillo Carlo Pellizzari di San Girolamo
Scuola Normale Superiore, p.zza dei Cavalieri 7, 56126 Pisa PI, Italy
F
Francesco Tosoni
Sant’Anna School of Advanced Studies, L’EMbeDS, p.zza Martiri della Libertà 33, 56127 Pisa PI, Italy