CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of extracting implicit engineering knowledge, such as algorithms and paradigms, from massive codebases using existing tools. We propose a large language model (LLM)-based approach for constructing a code knowledge graph. By leveraging specialized LLMs to extract code semantics and integrating SPARQL queries, Deep Research Agents, and hierarchical aggregation techniques, our method links extracted information to Wikidata through a three-stage pipeline to build an open-taxonomy code knowledge graph. Additionally, we introduce an LLM-assisted calibration protocol for quality assurance. Applied to 167 million code files, this framework yields CodeGraph, the first large-scale, open-taxonomy code knowledge graph comprising 158 million nodes and one billion edges. This work enables the systematic organization and representation of software engineering knowledge at an unprecedented scale.
📝 Abstract
Public software repositories, like GitHub and Software Heritage Archive, store billions of files, yet extracting their implicit engineering knowledge ---i.e., the algorithms they implement, the paradigms they follow, the patterns they instantiate, and the application domains they serve--- remains challenging, as current tools are constrained to syntactic and token-level analysis. We present a pipeline for building an open-taxonomy semantic annotation of source code using a code-specialised Large Language Model. The extracted entities are grounded in Wikidata through a three-stage linking procedure: a deterministic SPARQL stage handles unambiguous entities, a Deep Research Agent resolves the residual long tail, and a hierarchy-rollup stage imports the parent-of closure of each resolved Wikidata identifier. The resulting annotations are materialised as a source-code-specific open-taxonomy knowledge graph. We further introduce a calibrated quality-assurance protocol that quantifies annotation precision by combining a small human gold set with an LLM-as-a-judge filter. We applied our pipeline to the 167 million files of the Stack-Edu corpus, creating the first known large-scale open-taxonomy knowledge graph for source code. Our graph, named CodeGraph, contains approximately 158 million nodes, which include around 145 million files, about 63,000 extracted concept entities (such as algorithms, paradigms, design patterns, and application domains), and roughly 19,800 grounded Wikidata entities. Furthermore, CodeGraph features approximately 1 billion typed edges that connect files to their respective concepts, link these concepts to their grounded Wikidata identifiers, and relate them to their parent categories, covering 14 programming languages.
Problem

Research questions and friction points this paper is trying to address.

source code
knowledge graph
semantic annotation
software repositories
engineering knowledge
Innovation

Methods, ideas, or system contributions that make the work stand out.

Open-Taxonomy Knowledge Graph
Source Code Annotation
Wikidata Grounding
Large Language Model
Deep Research Agent