🤖 AI Summary
This study addresses the limitations of traditional bibliometric indicators, which overlook academic network topology and struggle to detect collusive research misconduct. The authors propose a novel anomaly detection framework that constructs a heterogeneous, multivariate graph from OpenAlex data, incorporating seven node and edge types. By integrating network projection, interpretable structural metrics, community detection, and three specialized anomaly pattern filters, the method generates a ranked list of suspicious entities accompanied by explicit structural evidence—without relying on binary classification. The approach uncovers metric confounding issues and demonstrates that graph-based prestige measures exhibit strong robustness against citation manipulation, outperforming conventional metrics by an order of magnitude in resilience. Experiments successfully reconstruct research teams and identify interdisciplinary bridges in VSB University data; journal-level analysis confirms disciplinary breadth as a reliable signal (AUC=0.70). An open-source tool, apnet, enables minute-scale analyses.
📝 Abstract
The academic publishing ecosystem is a vast, heterogeneous network of works, authors, institutions, journals, and topics. Traditional scientometrics reduces it to isolated tabular indicators (h-index, Impact Factor) that ignore topological context and are not designed to capture coordinated illegitimate practices. Building on our companion review, which proposed graph analysis of publishing integrity, this paper implements that approach. We define a heterogeneous multivariate graph model over OpenAlex open data (seven node types, seven edge types) and a methodology based on projections (citation and co-authorship networks), interpretable structural metrics, community detection, and three screening detectors of anomalous publishing patterns. We deliberately avoid binary classification: detectors return ranked candidates with explicit structural evidence for human assessment. On the institutional corpus of VSB - Technical University of Ostrava (2020-2025) with its one-hop citation neighbourhood, community detection recovers real research groups, centralities identify cross-disciplinary bridges, and the screenings flag dense co-authorship cliques, locally closed citation loops, and thematically isolated venues. On a second, venue-centric corpus with external ground truth (journals delisted by Scopus and DOAJ) and size-matched controls, a naive case-control design yields seemingly strong but spurious detectors (a prominence confound), whereas after matching the only robust signal is the breadth of disciplinary scope (AUC 0.70); an open graph-based prestige measure (PageRank over the journal citation network) tracks a JIF proxy while being an order of magnitude more resistant to citation gaming than count-based indicators. We release the method as the open-source library apnet with a reproducible CLI workflow and a web interface; the analysis runs on commodity hardware in minutes.