Presenting a classifier to detect research contributions in OpenAlex

📅 2025-07-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses bibliometric bias arising from the conflation of research articles with non-research content (e.g., editorials, abstracts, letters, prefaces) in the OpenAlex database. To mitigate this, we develop a lightweight, open-metadata–based machine learning classifier for binary document-type classification. Leveraging features including title, abstract text, citation patterns, and structured metadata fields, we train and optimize a supervised model to accurately distinguish non-research publications. Our key contribution is the first large-scale, systematic re-annotation of document types within an open citation dataset comprising over 4.27 million records. The optimized model achieves an F1-score of 0.95, enabling the identification and correction of 4.58 million non-research records—10.75% of the corpus—thereby substantially improving data purity and the reliability of scholarly analytics.

Technology Category

Data Mining & Knowledge Management: Linked Open Data, Knowledge Graphs & KB CompletionApplication Domains: Humanities & Computational Social ScienceNatural Language Processing: Information Extraction

Application Category

Web Mining and Content Analysis: Web data provenance, reliability, and authenticityEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsSemantics and Knowledge: Scalable techniques for the creation, curation, publication, maintenance, and consumption of large, Web-based, structured, reusable, knowledge graphs and ontologies
📝 Abstract
This paper introduces a document type classifier with the purpose to optimise the distinction between research and non-research journal publications in OpenAlex. Based on open metadata, the classifier can detect non-research or editorial content within a set of classified articles and reviews (e.g. paratexts, abstracts, editorials, letters). The classifier achieves an F1-score of 0,95, indicating a potential improvement in the data quality of bibliometric research in OpenAlex when applying the classifier on real data. In total, 4.589.967 out of 42.701.863 articles and reviews could be reclassified as non-research contributions by the classifier, representing a share of 10,75%
Problem

Research questions and friction points this paper is trying to address.

Classify research vs non-research publications in OpenAlex
Improve data quality in bibliometric research using metadata
Identify editorial content like letters and abstracts automatically
Innovation

Methods, ideas, or system contributions that make the work stand out.

Classifier distinguishes research from non-research publications
Uses open metadata to detect editorial content
Achieves high F1-score for improved data quality
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
N
Nick Haupka
Göttingen State and University Library, University of Göttingen