BioMedJImpact: A Comprehensive Dataset and LLM Pipeline for AI Engagement and Scientific Impact Analysis of Biomedical Journals

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic characterization of how artificial intelligence (AI) engagement influences the impact of biomedical journals. The authors construct BioMedJImpact, a dataset comprising 1.74 million PubMed Central articles, and introduce journal-level annual AI engagement rate as a novel quantitative metric. They propose a reproducible three-stage large language model (LLM) pipeline integrated with manual validation, collaboration network analysis, and bibliometric methods to enable content-aware scientometrics. Their findings reveal a positive association between team size and citation impact, and show that AI engagement rates significantly boosted journal impact factors only during 2020–2023. Manual evaluation further confirms that the LLM demonstrates consistent and reliable performance in identifying AI-related content and classifying subfields.
📝 Abstract
Assessing journal impact is central to scholarly communication, yet existing resources rarely capture how collaboration and artificial intelligence (AI) research jointly shape venue prestige in biomedicine. We present BioMedJImpact, a large-scale, biomedical-oriented dataset built from 1.74 million PubMed Central articles across 2,744 journals. BioMedJImpact integrates bibliometric indicators, collaboration features, and an LLM-derived AI engagement rate, defined as the proportion of AI-related articles within each journal-year. Specifically, AI engagement rate is extracted through a reproducible three-stage LLM pipeline. We analyze how collaboration intensity and AI engagement rate jointly influence scientific impact across two temporal subsets (2016-2019, 2020-2023). Two main patterns emerge: journals with larger author teams tend to have higher citation impact, while AI engagement rate is positively associated with Impact Factor only in the 2019 subset. To validate the LLM pipeline for deriving the AI engagement rate, we conduct human evaluation, confirming substantial agreement in AI relevance detection and consistent subfield classification. Together, BioMedJImpact provides both a comprehensive dataset at the interface of biomedicine and AI and a validated framework for scalable, content-aware scientometric analysis. Code and dataset are available at https://github.com/JonathanWry/BioMedJImpact.
Problem

Research questions and friction points this paper is trying to address.

journal impact
AI engagement
biomedical journals
scientific impact
collaboration
Innovation

Methods, ideas, or system contributions that make the work stand out.

large language model (LLM)
AI engagement rate
scientometric analysis
biomedical journals
reproducible pipeline
🔎 Similar Papers
No similar papers found.