Turning Citation Networks Inside Out: Studying Science Using Content-Based Knowledge Graphs from LLM-Derived Taxonomies

📅 2026-01-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional science mapping relies heavily on citations and metadata, which often fail to capture the intrinsic structure and evolution of scientific knowledge. This work proposes an “inside-out” paradigm for science mapping by leveraging large language models to construct domain-specific ontologies and encode scholarly papers into interpretable, content-level knowledge triples comprising “measurement method—data type—research question.” A normalized betweenness–connectivity ratio is introduced to identify structurally pivotal bridging components. Applied to a corpus of 617 studies on intergenerational wealth mobility, the approach reveals a stable backbone centered on regression methods and uncovers significant temporal patterns of component recombination, effectively illuminating the dynamic architecture of the field’s knowledge system.

Technology Category

Data Mining & Knowledge Management: Graph Mining, Social Network Analysis & CommunityApplication Domains: Humanities & Computational Social ScienceKnowledge Representation and Reasoning: Ontologies

Application Category

Semantics and Knowledge: Scalable techniques for the creation, curation, publication, maintenance, and consumption of large, Web-based, structured, reusable, knowledge graphs and ontologiesWeb Mining and Content Analysis: Bridging structured and unstructured dataGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Scientific fields are often mapped using citations and metadata, despite knowledge being transmitted primarily through content. We introduce an'inside-out'approach that reconstructs field structure directly from text by representing each paper as a small set of interpretable knowledge components. Using a large language model to induce domain-specific taxonomies and label papers, each publication is encoded as a triplet of measure, data type, and research-question type. These triplets define a knowledge graph with edges weighted by shared papers. Applied to 617 studies on intergenerational wealth mobility, the graph reveals a stable methodological backbone centered on regression-based mobility measures, alongside substantial temporal variation in component recombination. We further utilize normalized betweenness-to-connectivity ratios to identify components and pairings that act as structural bridges disproportionate to their prevalence. This content-derived, taxonomy-driven mapping complements citation-based approaches by exposing the evolving architecture of methods, data, and questions that define a field.
Problem

Research questions and friction points this paper is trying to address.

citation networks
knowledge graphs
scientific mapping
content-based analysis
intergenerational wealth mobility
Innovation

Methods, ideas, or system contributions that make the work stand out.

knowledge graph
large language model
taxonomy induction
content-based mapping
scientific field structure
S
Seorin Kim
Data Analytics Lab, Vrije Universiteit Brussel, 1050, Brussel, Belgium
V
Vincent Holst
Data Analytics Lab, Vrije Universiteit Brussel, 1050, Brussel, Belgium
Vincent Ginis
Vincent Ginis
Vrije Universiteit Brussel / Harvard University
Physics | Machine Learning