🤖 AI Summary
This study addresses bibliometric bias arising from the conflation of research articles with non-research content (e.g., editorials, abstracts, letters, prefaces) in the OpenAlex database. To mitigate this, we develop a lightweight, open-metadata–based machine learning classifier for binary document-type classification. Leveraging features including title, abstract text, citation patterns, and structured metadata fields, we train and optimize a supervised model to accurately distinguish non-research publications. Our key contribution is the first large-scale, systematic re-annotation of document types within an open citation dataset comprising over 4.27 million records. The optimized model achieves an F1-score of 0.95, enabling the identification and correction of 4.58 million non-research records—10.75% of the corpus—thereby substantially improving data purity and the reliability of scholarly analytics.
📝 Abstract
This paper introduces a document type classifier with the purpose to optimise the distinction between research and non-research journal publications in OpenAlex. Based on open metadata, the classifier can detect non-research or editorial content within a set of classified articles and reviews (e.g. paratexts, abstracts, editorials, letters). The classifier achieves an F1-score of 0,95, indicating a potential improvement in the data quality of bibliometric research in OpenAlex when applying the classifier on real data. In total, 4.589.967 out of 42.701.863 articles and reviews could be reclassified as non-research contributions by the classifier, representing a share of 10,75%