Score
Collecting, processing, and analyzing bibliographic and citation data to measure research production, impact, and concentration over time and across countries/institutions, and to identify methodological or resource gaps in scholarly literatures.
This study identifies a structural imbalance in interdisciplinary data sharing: high reuse rates in STEM fields contrast sharply with low adoption in humanities and social sciences, while persistent undercitation of datasets impedes evidence-based policy and infrastructure development. Leveraging full-text PubMed articles, we construct the first multidisciplinary dataset—enabling simultaneous identification of data mentions and classification of data-related intents—by integrating natural language processing, full-text pattern recognition, cross-disciplinary bibliometrics, and time-series modeling. Key findings include: (1) a marked acceleration in data publication post-2012; (2) highest data publishing activity in business/management and creative arts, yet highest reuse in biological and agricultural sciences; and (3) consistently low dataset citation rates, revealing critical bottlenecks in discoverability and format interoperability. These empirically grounded insights advance data governance frameworks and support formal recognition of datasets as independent scholarly outputs.
Institutions face significant challenges in systematically tracking their affiliated research data publications, primarily due to the widespread absence, inconsistency, or non-standardization of institutional attribution metadata (e.g., missing or ambiguous institutional names, lack of persistent identifiers such as DOIs) in existing data repositories. Method: We propose the first open-source, institution-centric workflow for tracking data publications, integrating over 70 open APIs and employing multi-source metadata harvesting, normalization, and automated aggregation—thereby reducing reliance on explicit attribution signals like DOIs or manually curated affiliations. Contribution/Results: The workflow enables efficient discovery and consolidation of over 4,000 cross-platform datasets. Evaluation demonstrates substantial improvements in institutional data discoverability, coverage breadth, and retrieval efficiency. This work delivers a reusable technical infrastructure to support research administration and data governance at institutional and systemic levels.
This study investigates the existence of “research silos” between academic scholars and academic librarians in Canada’s Library and Information Science (LIS) field. Methodologically, it constructs a novel, comprehensive Canadian LIS publication database—uniquely incorporating non-traditional scholarly outputs by academic librarians (e.g., practice reports, white papers, trade journals)—and applies bibliometric analysis, co-occurrence networks, citation coupling, and collaboration network mining to compare research output, thematic distribution, publication venues, and citation patterns between the two groups. Results reveal significant divergence in topical focus and venue selection, minimal collaborative authorship, and weak cross-group citation, confirming a structural research divide. By extending bibliometric analysis beyond peer-reviewed journal articles—a longstanding limitation in LIS scholarship—this work provides empirical evidence for bridging the theory–practice gap and advances methodological rigor in LIS research evaluation.
This work proposes a novel dataset discovery framework that leverages citation contexts from scientific papers to better capture the semantic intent behind research queries, addressing the limitations of existing dataset search engines that rely primarily on metadata and keyword matching and consequently suffer from low recall. By treating citation context as the core signal—combined with large-scale context extraction, large language model–guided pattern recognition, and provenance-preserving entity resolution—the approach significantly reduces dependence on incomplete or inconsistent metadata. Evaluated on eight computer science queries, the method achieves an average normalized recall of 47.47% (peaking at 81.82%), substantially outperforming Google Dataset Search and DataCite Commons. The framework’s novelty and practical utility have been affirmed by domain experts across multiple disciplines.
This study addresses the lack of macro-level evolutionary mapping in the interdisciplinary field of library and cultural studies. Employing bibliometric analysis and visualization tools (CiteSpace, VOSviewer), it systematically examines 890 publications from Web of Science Core Collection (2010–2019) to analyze growth trends, collaborative patterns across authors, institutions, and countries, and knowledge diffusion dynamics. Results reveal the United States as the most prolific country (244 papers), Wuhan University’s School of Information Management as the leading institution, and *The Journal of Academic Librarianship* as the dominant journal. International collaboration exhibits a pronounced imbalance, centered on North America and Europe, with emerging participation from Asia. This is the first large-scale scientometric investigation of this interdisciplinary domain, providing empirical evidence for strategic discipline planning and cross-domain research collaboration.
This study investigates how academic age influences methodological choices among scholars in Library and Information Science (LIS). Drawing on a corpus of 26,677 articles published between 1990 and 2023 in 14 leading journals, the authors employ author disambiguation techniques to compute academic age and introduce, for the first time, the CogFT model to automatically classify research methods, complemented by Top2Vec for topic modeling. The analysis reveals a sustained decline in the use of theoretical approaches alongside significant increases in experimental and bibliometric methods. Methodological diversity peaks among mid-career researchers and is lowest among late-career scholars. These findings illuminate evolving patterns in research methodology within LIS and offer empirical guidance for early-career researchers in selecting appropriate methods.
This work addresses the lack of effective monitoring of dataset usage in scholarly literature, which undermines citation transparency, impact traceability, and reproducibility. To tackle this challenge, the study introduces the first application of the multi-task GLiNER framework to dataset usage monitoring, jointly performing dataset mention extraction, relation identification, and usage context classification. The approach integrates synthetic data generation with a large language model (LLM)-driven re-verification mechanism to mitigate issues of annotation scarcity and ambiguous citations. This combination significantly enhances the accuracy, coverage, and label consistency of dataset mention detection, enabling end-to-end, unconstrained tracking of data citations across diverse scientific texts and advancing the development of open-source tools for scholarly data provenance.
This study addresses the longstanding challenge of effectively measuring the real-world impact of research data reuse in open science. For the first time, it leverages large language models (LLMs) and generative artificial intelligence to conduct large-scale analysis of scholarly texts, automatically identifying instances where published works reuse existing datasets. The proposed approach achieves substantially higher detection accuracy compared to conventional methods, revealing a data reuse rate of 43%—far exceeding estimates derived from traditional bibliometric indicators. These findings demonstrate that current metrics systematically underestimate the benefits of data sharing and establish a novel paradigm for evaluating the impact of open science initiatives.
This work proposes BIP! Scholar, a novel researcher profiling platform that addresses the limitations of existing systems, which predominantly focus on publications and bibliometric indicators and thus fail to capture the full spectrum of scholarly contributions in context. BIP! Scholar introduces a template-driven architecture that enables researchers to flexibly construct structured profiles based on trajectories, narratives, or hybrid models, tailored to specific evaluation or presentation needs. By integrating customizable templates, multi-faceted research activity modeling, and a user-configurable interface, the platform dynamically represents non-traditional outputs, roles, and scholarly activities. This approach significantly enhances the contextual adaptability and evaluative utility of researcher profiles and empowers assessment experts to design and test new profiling templates.
This study investigates the divergent patterns and asynchronous evolution of research methodology adoption across countries in the field of Library and Information Science (LIS) from 1990 to 2019. By integrating manual annotation with deep learning models, the authors present the first large-scale cross-national classification of methodologies employed in 5,281 international journal articles from 81 countries. The findings reveal that national methodological practices exhibit distinct and statistically significant distributional patterns, yet these differences have gradually diminished over time at both national and international levels. This work offers novel methodological insights and empirical evidence for understanding country-specific trajectories in LIS development and fostering international scholarly collaboration.