Score
Building and analyzing networks of scholarly artifacts and contributors (citations, co-authorship, collaboration links) using centrality and related measures to quantify impact, disruptiveness, and translational signals (e.g., patent-paper links) for search and evaluation tasks.
This study investigates structural and evolutionary differences in author collaboration networks between data mining and software engineering, aiming to uncover community organization, temporal dynamics, and distributions of influential scholars and institutions. Method: Leveraging Google Scholar data from 2000–2021, we construct and comparatively analyze the two networks using graph-theoretic metrics—including sparsity, clustering coefficient, small-worldness, modularity, and degree centrality—alongside temporal modeling and Louvain community detection. Contribution/Results: We empirically demonstrate that both fields exhibit highly clustered, localized small-world communities; however, their influence distributions diverge markedly: data mining displays a hub-and-spoke pattern dominated by a few highly connected authors, whereas software engineering manifests a multi-centric, diffused structure. The analysis identifies core author clusters and high-output institutions in each domain, offering quantitative evidence to inform interdisciplinary collaboration strategies and disciplinary development assessment.
Citation disparities in AI top-tier conferences persist despite controlling for paper quality, suggesting unobserved structural drivers. Method: Leveraging 17,942 papers from NeurIPS, ICML, and ICLR (2005–2024), we propose Harmonic Closeness Temporal Centrality with Decay (HCTCD)—a temporal, collaboration-strength-weighted centrality measure—and introduce Beta regression to model citation percentile ranks. Contribution/Results: We demonstrate that team-level exponentially weighted centrality aggregation substantially outperforms individual- or rank-based aggregation; long-term centrality exerts significantly greater influence than short-term metrics. Integrating HCTCD reduces mean squared error in citation prediction by 2.4%–4.8%. This work provides the first systematic empirical evidence that network structural bias—rather than content quality alone—dominates citation distribution in AI research, offering both novel interpretability and a quantifiable, fairness-aware tool for scholarly evaluation.
To address the challenge of early identification of scientific innovation, this paper proposes a forward-looking prediction framework for assessing the future academic impact of unpublished research directions. Methodologically, we construct a dynamic knowledge graph comprising 21 million scholarly papers, integrating semantic content and citation relationships—enabling, for the first time, quantitative prediction of impact for pre-publication scientific ideas. We introduce a joint modeling paradigm combining temporal graph neural networks, multimodal semantic embeddings, and dynamic link prediction. Our approach achieves an AUC exceeding 0.91, significantly outperforming baseline methods in identifying high-impact emerging topics at early stages. This work delivers an interpretable and scalable core predictive capability for AI-driven scientific discovery (“Artificial Muse”), thereby facilitating interdisciplinary innovation and optimizing the allocation of research resources.
Existing citation dynamics models are domain-specific and mechanistically fragmented, failing to explain cross-domain phenomena such as “delayed recognition” and the ubiquity of “sleeping beauties.” Method: We propose the first unified model integrating three fundamental mechanisms—cumulative advantage, temporal decay, and structural embedding—and validate it via multi-source citation network mining, temporal modeling, and cross-domain comparative analysis across scientific publications, legal cases, and patents. Contribution/Results: We empirically establish the cross-domain universality of skewed citation distributions—including sleeping beauties—across all three knowledge systems. Our model significantly outperforms state-of-the-art baselines in reproducing empirical citation evolution and achieves 12–28% higher accuracy in predicting high-impact nodes. This work reveals universal principles governing citation dynamics beyond disciplinary boundaries and provides a generalizable framework for modeling knowledge diffusion and impact accumulation.
This paper addresses the limitations of text-based approaches for interdisciplinary literature identification—namely, high computational cost and poor interpretability—by proposing a purely network-structural method. It models citation networks as directed acyclic graphs (DAGs) and introduces “diversity centrality,” a novel metric that integrates transitive reduction with degree centrality to identify pivotal papers bridging densely connected, multi-disciplinary subgroups. By applying topological reduction to eliminate redundant transitive paths, the method accentuates cross-domain hub papers. Experiments across multiple real-world citation networks demonstrate that the approach achieves interdisciplinary impact detection performance comparable to state-of-the-art text-analytic methods, while being computationally efficient, parameter-free, and fully interpretable. It thus establishes a scalable, transparent, and structurally grounded paradigm for assessing interdisciplinary research.
This study addresses the overreliance on citation counts in traditional research evaluation, which overlooks the intermediate pathways of knowledge dissemination. It introduces “citation pathways” as a novel dimension in scientometrics and formally defines two key intermediary structures within them: Interpretive Knowledge Nodes (IKNs) and Citation Compression Layers (CCLs). By integrating normative citation structure analysis, thought experiments, and a simplified “citation gravity” model, the work reveals how artificial intelligence reshapes the production costs of citable knowledge intermediaries and alters the evolutionary dynamics of citation networks. The findings demonstrate that, under compliant citation practices, the positional effects of entities within these pathways significantly influence the validity of impact assessments, highlighting potential misalignments in institutional incentives under extreme conditions and thereby redefining the boundaries of academic impact measurement.
Traditional citation networks treat all references uniformly, making it difficult to identify the core sources that genuinely inspire a study and thereby compromising the accuracy of impact assessment. This work proposes a novel approach that systematically leverages large language models (LLMs) with two prompting strategies to automatically detect seminal citations from full-text articles, constructing a backbone citation network that captures the essential structure of scientific knowledge. Analyses reveal that, although smaller in scale, this backbone network exhibits non-random topology with higher heterogeneity in in-degree distribution. Its topological properties—such as modularity, transitivity, and degree assortativity—systematically differ from those of the full citation network. Nevertheless, rankings of highly cited papers and authors show strong consistency between the two networks, suggesting that despite containing redundancy, the full network remains effective in reflecting relative scholarly influence.
Existing approaches struggle to model the dynamic co-evolution among authors, references, and keywords in large-scale, fine-grained scientific collaboration data and cannot effectively test multiple competing hypotheses about the mechanisms of collective knowledge production. This work extends the Relational Hyper-Event Model (RHEM) to dynamic tripartite hypergraphs, introducing a unified framework that captures the joint generative process of multidimensional, heterogeneous entities within hyper-events of arbitrary size while explicitly controlling intra- and inter-set dependencies. The proposed method enables simultaneous modeling of complex interactions and rigorous statistical hypothesis testing on real-world scholarly data, offering an interpretable and quantifiable comparison of the drivers underlying collective knowledge creation.
This study addresses the limitation of existing research that relies predominantly on macro-level indicators to assess knowledge proximity between academia and industry, which often fails to capture fine-grained knowledge linkages. To overcome this, the authors propose a multidimensional quantification framework integrating entity-based and semantic approaches. Specifically, they leverage pre-trained language models to extract fine-grained knowledge entities from scholarly texts and analyze their sequential overlap and network topological features. Additionally, an unsupervised contrastive learning method is introduced to evaluate the convergence of semantic spaces across institutional boundaries, complemented by citation distribution analysis to uncover the relationship between bidirectional knowledge flows and similarity. This work presents the first integration of fine-grained entity recognition with semantic contrastive learning, empirically demonstrating that technological transitions significantly enhance academic–industrial knowledge proximity, offering textual evidence of their co-evolution and revealing a weakening of academic dominance during paradigm shifts.
This study investigates how individual scientific and technological productivity is shaped by collaborators’ behavior and network position, and clarifies whether scientific and technological collaboration act as complements or substitutes. Leveraging a simultaneous-equation network framework, the authors construct a two-layer collaboration network integrating co-authored papers and co-invented patents. Identification relies on a modified Katz-Bonacich centrality measure, an instrumental variable approach exploiting exogenous dyadic attributes to predict link formation, and community fixed effects. The analysis provides the first joint dynamic evidence of peer effects in both science and invention, revealing an asymmetric mechanism through which science drives innovation: network centrality significantly enhances both types of output, and scientific productivity positively spurs technological output, whereas the reverse effect is statistically insignificant. These findings offer a microfoundation for policies promoting synergistic innovation.