Score
Designs and analyzes quantitative maps of scholarly knowledge by extracting and modeling bibliometric relationships (e.g., co-citation, co-occurrence, citation and venue linkages) to produce network or landscape representations. Builds cluster and visualization outputs that identify disciplinary clusters and bridges, central bodies of work, and principal journals or venues.
Faced with the challenge of analyzing evolving research landscapes amid exponential growth in scientific literature, this paper proposes LitLA—a novel end-to-end literature analysis workflow. LitLA pioneers a full-lifecycle knowledge graph construction paradigm tailored to the scholarly ecosystem. It integrates publication metadata with advanced techniques including knowledge graph construction, temporal graph embedding, dynamic network analysis, and interpretable topic modeling to enable multidimensional modeling and evolutionary inference of the MOEA/D research domain. The resulting knowledge graph encompasses over 5,400 papers, 10,000 authors, 1,600 institutions, and 78,000 keywords. It supports fine-grained tracking of thematic evolution, visualization of academic community dynamics, and interpretable forecasting of future research trends. By unifying heterogeneous scholarly signals within a scalable, graph-based framework, LitLA significantly enhances the systematicity and scalability of large-scale intelligent literature analysis.
Novice learners in network science lack systematic, interdisciplinary guidance. Method: We propose the “Network Scientist Training Map” framework, integrating diverse methodological tools—including graph theory, random graph models, spectral graph theory, network embedding, statistical inference, dynamical modeling, and interpretable machine learning—restructuring core paradigms in an accessible, unified manner that emphasizes methodological synthesis over technical accumulation. Contribution/Results: First, we formally establish pure network science as an autonomous academic discipline. Second, we construct a structured, full-stack knowledge graph covering foundational concepts and techniques. Third, by transcending the limitations of single-discipline textbooks, the framework enables learners with no prior background to rapidly develop a coherent, systems-level conceptual framework—thereby advancing network science from an application-oriented auxiliary field toward ontological independence.
Traditional science mapping relies heavily on citations and metadata, which often fail to capture the intrinsic structure and evolution of scientific knowledge. This work proposes an “inside-out” paradigm for science mapping by leveraging large language models to construct domain-specific ontologies and encode scholarly papers into interpretable, content-level knowledge triples comprising “measurement method—data type—research question.” A normalized betweenness–connectivity ratio is introduced to identify structurally pivotal bridging components. Applied to a corpus of 617 studies on intergenerational wealth mobility, the approach reveals a stable backbone centered on regression methods and uncovers significant temporal patterns of component recombination, effectively illuminating the dynamic architecture of the field’s knowledge system.
Traditional scientometric approaches—topic modeling (TM, e.g., LDA) and citation clustering (CC, e.g., VOS clustering)—exhibit divergent assumptions about scientific structure, yet their comparative representational capacities in mapping cardiovascular science remain underexplored. Method: This study systematically compares TM and CC across three dimensions—disciplinary architecture, responsiveness to societal needs, and delineation of academic micro-communities—and introduces a cross-method mapping framework to align topics and clusters. Contribution/Results: Empirical analysis reveals only weak correspondence (<33% document overlap) between TM-derived topics and CC-derived clusters. TM proves more sensitive to policy-relevant societal themes (e.g., public health interventions), whereas CC more accurately reconstructs knowledge evolution trajectories and disease-subtype-driven scholarly communities. The findings demonstrate functional complementarity: neither method alone suffices to capture the multidimensional nature of scientific fields. This work provides a methodological foundation for informed tool selection in scientometrics and science policy research.
In citation networks, temporal constraints induce asymmetry in the adjacency matrix (with missing lower-triangular entries), leading to unidentifiable common factors in standard factorization models. Method: We propose a dual-space common-factor model that separately encodes citing and cited behaviors of papers, establishing the first identifiable common-factor decomposition framework for asymmetric, time-constrained adjacency matrices, and designing a memory-efficient, time-aware matrix completion algorithm. Contribution/Results: Evaluated on the largest statistical literature dataset to date—256,000 papers spanning 1898–2024—we achieve joint embedding and topic modeling for high-dimensional sparse citation networks. Our approach uncovers 11 semantically coherent and interpretable subfield common-factor structures, advancing temporal network representation learning and scientometric analysis with a novel, theoretically grounded paradigm.
This study addresses the limitations of existing academic platforms in providing fine-grained author metadata and geographic visualization, which are either unavailable or prohibitively costly. The authors propose an automated analytical framework that leverages a single Google Scholar user ID and integrates data from five sources—Google Scholar, OpenAlex, CrossRef, Semantic Scholar, and OpenStreetMap—through a five-stage pipeline for paper parsing, author disambiguation, and geocoding. Key innovations include a Unicode-resilient metadata parser, a two-stage institutional similarity–based disambiguation mechanism, and a city-level location repair method using OpenAlex. The system increases city-level geographic coverage for authors from 0% to approximately 60%, reduces h-index attribution errors by up to ninefold, and produces interactive HTML maps alongside structured analytical reports.
This study addresses the overreliance on citation counts in traditional research evaluation, which overlooks the intermediate pathways of knowledge dissemination. It introduces “citation pathways” as a novel dimension in scientometrics and formally defines two key intermediary structures within them: Interpretive Knowledge Nodes (IKNs) and Citation Compression Layers (CCLs). By integrating normative citation structure analysis, thought experiments, and a simplified “citation gravity” model, the work reveals how artificial intelligence reshapes the production costs of citable knowledge intermediaries and alters the evolutionary dynamics of citation networks. The findings demonstrate that, under compliant citation practices, the positional effects of entities within these pathways significantly influence the validity of impact assessments, highlighting potential misalignments in institutional incentives under extreme conditions and thereby redefining the boundaries of academic impact measurement.
Traditional citation networks treat all references uniformly, making it difficult to identify the core sources that genuinely inspire a study and thereby compromising the accuracy of impact assessment. This work proposes a novel approach that systematically leverages large language models (LLMs) with two prompting strategies to automatically detect seminal citations from full-text articles, constructing a backbone citation network that captures the essential structure of scientific knowledge. Analyses reveal that, although smaller in scale, this backbone network exhibits non-random topology with higher heterogeneity in in-degree distribution. Its topological properties—such as modularity, transitivity, and degree assortativity—systematically differ from those of the full citation network. Nevertheless, rankings of highly cited papers and authors show strong consistency between the two networks, suggesting that despite containing redundancy, the full network remains effective in reflecting relative scholarly influence.
This work addresses critical limitations of traditional science, technology, and innovation (STI) analysis—namely its reliance on static indicators that suffer from time lags, shallow semantic representation, and an inability to capture the nonlinear dynamics of knowledge ecosystems. To overcome these challenges, the study proposes a validation-centric, five-layer hybrid framework that constructs a dynamic, versioned knowledge graph from open scholarly data and integrates constrained large language models (LLMs) for structured semantic enrichment. Reliability is ensured through a multi-tiered validation pipeline combining structural, evidential, comparative, and expert-based verification, augmented by a traceability mechanism. This approach significantly enhances the semantic depth, timeliness, and credibility of STI analysis while upholding scientific evidentiary standards, thereby enabling robust detection of emerging trends, mapping of technology transfer pathways, and policy-relevant gap analysis.
This study investigates whether low-impact journals exhibit denser and more reciprocal author citation patterns and the resulting distortions in bibliometric indicators. Leveraging Crossref data, journals are stratified by discipline-normalized Eigenfactor percentiles to distinguish low-impact (Case) from high-impact (Control) groups, with author-matching controls employed to compare citation network structures. The analysis reveals that low-impact journals form insular “citation economies,” manifesting a pronounced bifurcation into “two worlds.” It further identifies 277 high-purity anomalous citation clusters characterized by core–periphery topologies. Findings indicate that authors in low-impact journals display 6.7 times higher mutual citation rates, 4.7 times greater reciprocity, and an 11-fold increase in citation clique strength, confirming the presence of directed citation flows rather than equitable mutual referencing.