scholarly network analysis

Building and analyzing networks of scholarly artifacts and contributors (citations, co-authorship, collaboration links) using centrality and related measures to quantify impact, disruptiveness, and translational signals (e.g., patent-paper links) for search and evaluation tasks.

scholarlynetworkanalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Examining Different Research Communities: Authorship Network

Aug 24, 2024
SG
Shrabani Ghosh
🏛️ University of North Carolina at Charlotte

This study investigates structural and evolutionary differences in author collaboration networks between data mining and software engineering, aiming to uncover community organization, temporal dynamics, and distributions of influential scholars and institutions. Method: Leveraging Google Scholar data from 2000–2021, we construct and comparatively analyze the two networks using graph-theoretic metrics—including sparsity, clustering coefficient, small-worldness, modularity, and degree centrality—alongside temporal modeling and Louvain community detection. Contribution/Results: We empirically demonstrate that both fields exhibit highly clustered, localized small-world communities; however, their influence distributions diverge markedly: data mining displays a hub-and-spoke pattern dominated by a few highly connected authors, whereas software engineering manifests a multi-centric, diffused structure. The analysis identifies core author clusters and high-output institutions in each domain, offering quantitative evidence to inform interdisciplinary collaboration strategies and disciplinary development assessment.

Analyzing co-authorship networks in Data Mining and Software EngineeringExamining publication trends and network structures across disciplinesIdentifying influential authors and organizations within research communities

Beyond Content: How Author Network Centrality Drives Citation Disparities in Top AI Conferences

Dec 25, 2025
RJ
Renlong Jie
🏛️ Northwestern Polytechnical University | Yunnan University of Finance and Economics

Citation disparities in AI top-tier conferences persist despite controlling for paper quality, suggesting unobserved structural drivers. Method: Leveraging 17,942 papers from NeurIPS, ICML, and ICLR (2005–2024), we propose Harmonic Closeness Temporal Centrality with Decay (HCTCD)—a temporal, collaboration-strength-weighted centrality measure—and introduce Beta regression to model citation percentile ranks. Contribution/Results: We demonstrate that team-level exponentially weighted centrality aggregation substantially outperforms individual- or rank-based aggregation; long-term centrality exerts significantly greater influence than short-term metrics. Integrating HCTCD reduces mean squared error in citation prediction by 2.4%–4.8%. This work provides the first systematic empirical evidence that network structural bias—rather than content quality alone—dominates citation distribution in AI research, offering both novel interpretability and a quantifiable, fairness-aware tool for scholarly evaluation.

Examining how author network centrality influences citation disparities in AI conferencesProposing network-aware assessment to reduce structural biases in scientific recognitionQuantifying the impact of long-term centrality on citation percentiles using novel metrics

Forecasting high-impact research topics via machine learning on evolving knowledge graphs

Feb 13, 2024
XG
Xuemei Gu
🏛️ Max Planck Institute for the Science of Light

To address the challenge of early identification of scientific innovation, this paper proposes a forward-looking prediction framework for assessing the future academic impact of unpublished research directions. Methodologically, we construct a dynamic knowledge graph comprising 21 million scholarly papers, integrating semantic content and citation relationships—enabling, for the first time, quantitative prediction of impact for pre-publication scientific ideas. We introduce a joint modeling paradigm combining temporal graph neural networks, multimodal semantic embeddings, and dynamic link prediction. Our approach achieves an AUC exceeding 0.91, significantly outperforming baseline methods in identifying high-impact emerging topics at early stages. This work delivers an interpretable and scalable core predictive capability for AI-driven scientific discovery (“Artificial Muse”), thereby facilitating interdisciplinary innovation and optimizing the allocation of research resources.

Interdisciplinary InnovationPredictive AnalyticsResearch Impact

Uncovering the universal dynamics of citation systems: From science of science to law of law and patterns of patents

Jan 26, 2025
SK
Sadamori Kojaku
🏛️ Binghamton University | Indiana University | Massachusetts Institute of Technology | Harvard University | Southern University of Science and Technology | Northeastern University

Existing citation dynamics models are domain-specific and mechanistically fragmented, failing to explain cross-domain phenomena such as “delayed recognition” and the ubiquity of “sleeping beauties.” Method: We propose the first unified model integrating three fundamental mechanisms—cumulative advantage, temporal decay, and structural embedding—and validate it via multi-source citation network mining, temporal modeling, and cross-domain comparative analysis across scientific publications, legal cases, and patents. Contribution/Results: We empirically establish the cross-domain universality of skewed citation distributions—including sleeping beauties—across all three knowledge systems. Our model significantly outperforms state-of-the-art baselines in reproducing empirical citation evolution and achieves 12–28% higher accuracy in predicting high-impact nodes. This work reveals universal principles governing citation dynamics beyond disciplinary boundaries and provides a generalizable framework for modeling knowledge diffusion and impact accumulation.

Citation PatternsDelayed RecognitionUniversal Model

Diversity from the Topology of Citation Networks

Feb 16, 2018
VV
Vaiva Vasiliauskaite
🏛️ Imperial College London | Institute for Biomedical Engineering

This paper addresses the limitations of text-based approaches for interdisciplinary literature identification—namely, high computational cost and poor interpretability—by proposing a purely network-structural method. It models citation networks as directed acyclic graphs (DAGs) and introduces “diversity centrality,” a novel metric that integrates transitive reduction with degree centrality to identify pivotal papers bridging densely connected, multi-disciplinary subgroups. By applying topological reduction to eliminate redundant transitive paths, the method accentuates cross-domain hub papers. Experiments across multiple real-world citation networks demonstrate that the approach achieves interdisciplinary impact detection performance comparable to state-of-the-art text-analytic methods, while being computationally efficient, parameter-free, and fully interpretable. It thus establishes a scalable, transparent, and structurally grounded paradigm for assessing interdisciplinary research.

Assess interdisciplinary impact via citation network reductionClassify documents as interdisciplinary or intradisciplinaryValidate hypothesis using academic, legal, and patent citations

Latest Papers

What's happening recently
View more

This study addresses the overreliance on citation counts in traditional research evaluation, which overlooks the intermediate pathways of knowledge dissemination. It introduces “citation pathways” as a novel dimension in scientometrics and formally defines two key intermediary structures within them: Interpretive Knowledge Nodes (IKNs) and Citation Compression Layers (CCLs). By integrating normative citation structure analysis, thought experiments, and a simplified “citation gravity” model, the work reveals how artificial intelligence reshapes the production costs of citable knowledge intermediaries and alters the evolutionary dynamics of citation networks. The findings demonstrate that, under compliant citation practices, the positional effects of entities within these pathways significantly influence the validity of impact assessments, highlighting potential misalignments in institutional incentives under extreme conditions and thereby redefining the boundaries of academic impact measurement.

citation networkscitation pathwayknowledge intermediaries

Traditional citation networks treat all references uniformly, making it difficult to identify the core sources that genuinely inspire a study and thereby compromising the accuracy of impact assessment. This work proposes a novel approach that systematically leverages large language models (LLMs) with two prompting strategies to automatically detect seminal citations from full-text articles, constructing a backbone citation network that captures the essential structure of scientific knowledge. Analyses reveal that, although smaller in scale, this backbone network exhibits non-random topology with higher heterogeneity in in-degree distribution. Its topological properties—such as modularity, transitivity, and degree assortativity—systematically differ from those of the full citation network. Nevertheless, rankings of highly cited papers and authors show strong consistency between the two networks, suggesting that despite containing redundancy, the full network remains effective in reflecting relative scholarly influence.

backbone of sciencecitation networkscitation relevance

Existing approaches struggle to model the dynamic co-evolution among authors, references, and keywords in large-scale, fine-grained scientific collaboration data and cannot effectively test multiple competing hypotheses about the mechanisms of collective knowledge production. This work extends the Relational Hyper-Event Model (RHEM) to dynamic tripartite hypergraphs, introducing a unified framework that captures the joint generative process of multidimensional, heterogeneous entities within hyper-events of arbitrary size while explicitly controlling intra- and inter-set dependencies. The proposed method enables simultaneous modeling of complex interactions and rigorous statistical hypothesis testing on real-world scholarly data, offering an interpretable and quantifiable comparison of the drivers underlying collective knowledge creation.

collective productionmultipartite social networksrelational hyperevent models

This study addresses the limitation of existing research that relies predominantly on macro-level indicators to assess knowledge proximity between academia and industry, which often fails to capture fine-grained knowledge linkages. To overcome this, the authors propose a multidimensional quantification framework integrating entity-based and semantic approaches. Specifically, they leverage pre-trained language models to extract fine-grained knowledge entities from scholarly texts and analyze their sequential overlap and network topological features. Additionally, an unsupervised contrastive learning method is introduced to evaluate the convergence of semantic spaces across institutional boundaries, complemented by citation distribution analysis to uncover the relationship between bidirectional knowledge flows and similarity. This work presents the first integration of fine-grained entity recognition with semantic contrastive learning, empirically demonstrating that technological transitions significantly enhance academic–industrial knowledge proximity, offering textual evidence of their co-evolution and revealing a weakening of academic dominance during paradigm shifts.

academic-industry collaborationfine-grained knowledgeinstitutional divergence

This study investigates how individual scientific and technological productivity is shaped by collaborators’ behavior and network position, and clarifies whether scientific and technological collaboration act as complements or substitutes. Leveraging a simultaneous-equation network framework, the authors construct a two-layer collaboration network integrating co-authored papers and co-invented patents. Identification relies on a modified Katz-Bonacich centrality measure, an instrumental variable approach exploiting exogenous dyadic attributes to predict link formation, and community fixed effects. The analysis provides the first joint dynamic evidence of peer effects in both science and invention, revealing an asymmetric mechanism through which science drives innovation: network centrality significantly enhances both types of output, and scientific productivity positively spurs technological output, whereas the reverse effect is statistically insignificant. These findings offer a microfoundation for policies promoting synergistic innovation.

collaborationnetwork centralitypeer effects

Hot Scholars

EF

Emilio Ferrara

Professor of Computer Science at the University of Southern California
Human-Centered AISocial ComputingNetwork ScienceAI Safety
PS

Philip S. Yu

Professor of Computer Science, University of Illinons at Chicago
Data miningDatabasePrivacy
CZ

Chengzhi Zhang

Nanjing University of Science and Technology
Text MiningNatural Language ProcessingScience of Science
LL

Luca Luceri

Research Assistant Professor @University of Southern California - Information Sciences Institute
Computational Social ScienceNetwork ScienceMachine LearningSocial Media Manipulation
RH

Rong-Hua Li

Beijing Institute of Technology
Algorithms for (big) graphmatrixand sequence data