Score
Designs and builds systems that ingest full‑text scholarly documents and automatically assign article‑level labels for research methods and topics, including taxonomy‑based topic mappings (e.g., OpenAlex) and fine‑grained method categories. This work involves supervised text‑mining and classifier development, entity/method extraction and mapping pipelines, and evaluation and scaling strategies to produce coded method metadata, detect emerging topics, and support longitudinal or confound‑controlled analyses.
This study addresses the limitations of existing automatic multi-label classification approaches for scholarly papers, which predominantly rely on titles and abstracts—often insufficient for accurate labeling—while full texts are lengthy and exhibit uneven information distribution. To overcome this, the authors propose segmenting full-text articles by physical position and systematically evaluating the discriminative power of individual sections and their combinations. Their analysis reveals that middle-to-late and concluding sections carry higher informational value. By integrating these informative segments with bibliographic metadata, they construct an enhanced multi-label classification model. Experiments on a corpus of 1,954 library and information science journal articles demonstrate that the proposed cross-segment combination and metadata fusion strategy significantly improves classification accuracy, offering a novel and effective pathway for fine-grained method identification in academic texts.
This study addresses bibliometric bias arising from the conflation of research articles with non-research content (e.g., editorials, abstracts, letters, prefaces) in the OpenAlex database. To mitigate this, we develop a lightweight, open-metadata–based machine learning classifier for binary document-type classification. Leveraging features including title, abstract text, citation patterns, and structured metadata fields, we train and optimize a supervised model to accurately distinguish non-research publications. Our key contribution is the first large-scale, systematic re-annotation of document types within an open citation dataset comprising over 4.27 million records. The optimized model achieves an F1-score of 0.95, enabling the identification and correction of 4.58 million non-research records—10.75% of the corpus—thereby substantially improving data purity and the reliability of scholarly analytics.
Existing scientific classification systems struggle to capture fine-grained conceptual structures in research, while author-provided keywords—though specific—are often fragmented, redundant, and inconsistent in terminology. This work proposes a novel approach that integrates scientific text embeddings, large language models (LLMs), and graph-based community detection to construct an interpretable Concept layer beneath the Topic hierarchy of OpenAlex. By aggregating semantically related keywords into unified, reusable conceptual units and situating them within disciplinary hierarchies, this method enables a systematic transformation from heterogeneous terms to structured knowledge entities. It represents the first large-scale academic taxonomy to incorporate an LLM- and embedding-driven intermediate conceptual layer, substantially enhancing the precision and scalability of scholarly document organization, science mapping, research trend monitoring, and ontology construction.
Taxonomy construction for classifying unstructured text (e.g., personal goal statements) is time-intensive, prone to researcher bias, and suffers from poor reproducibility. Method: This paper proposes a human–AI collaborative, iterative text analysis paradigm that integrates top-down and bottom-up strategies to enable dynamic taxonomy generation, evaluation, refinement, and validation. Leveraging prompt engineering, it facilitates multi-turn collaboration between domain researchers and large language models (LLMs), with human feedback driving iterative taxonomy optimization. Intercoder reliability is quantified using Cohen’s κ within a structured coding framework. Results: Empirical evaluation in a life-domain dataset achieves κ > 0.85, significantly enhancing analytical efficiency, reliability, and reproducibility. This work pioneers deep integration of LLMs into the qualitative analysis closed loop, offering a novel methodology for low-bias, high-fidelity open-text classification.
This work proposes a novel dataset discovery framework that leverages citation contexts from scientific papers to better capture the semantic intent behind research queries, addressing the limitations of existing dataset search engines that rely primarily on metadata and keyword matching and consequently suffer from low recall. By treating citation context as the core signal—combined with large-scale context extraction, large language model–guided pattern recognition, and provenance-preserving entity resolution—the approach significantly reduces dependence on incomplete or inconsistent metadata. Evaluated on eight computer science queries, the method achieves an average normalized recall of 47.47% (peaking at 81.82%), substantially outperforming Google Dataset Search and DataCite Commons. The framework’s novelty and practical utility have been affirmed by domain experts across multiple disciplines.
This work addresses the lack of effective monitoring of dataset usage in scholarly literature, which undermines citation transparency, impact traceability, and reproducibility. To tackle this challenge, the study introduces the first application of the multi-task GLiNER framework to dataset usage monitoring, jointly performing dataset mention extraction, relation identification, and usage context classification. The approach integrates synthetic data generation with a large language model (LLM)-driven re-verification mechanism to mitigate issues of annotation scarcity and ambiguous citations. This combination significantly enhances the accuracy, coverage, and label consistency of dataset mention detection, enabling end-to-end, unconstrained tracking of data citations across diverse scientific texts and advancing the development of open-source tools for scholarly data provenance.
This study addresses key challenges faced by social science researchers when using large language models (LLMs) for text annotation—namely, poor reproducibility, annotation errors that compromise statistical inference, and high technical barriers. To overcome these issues, the authors propose the first end-to-end LLM-based text annotation framework tailored specifically for the social sciences and humanities (SSH). The framework integrates structured prompt engineering, open-source LLM API integration, cross-validation, and error propagation modeling, with an explicit emphasis on avoiding prompt overfitting and quantifying annotation uncertainty. Implemented in both Python and R, this approach establishes a transparent, reproducible, and scalable workflow that substantially enhances the reliability, efficiency, and methodological rigor of automated text annotation in SSH research.
This study addresses the current lack of interdisciplinary understanding regarding the integration pathways, efficacy boundaries, and systemic risks of large language models (LLMs) across natural sciences, social sciences, and humanities. Through a systematic literature review and illustrative case analyses, it critically evaluates the deployment of LLMs throughout the research lifecycle—including hypothesis generation, literature synthesis, data analysis, and scholarly writing. The work identifies ten previously underappreciated systemic risks, such as diminished researcher autonomy, AI-induced confirmation bias, ambiguous authorship, and inequitable access to technology. It further demonstrates how LLMs, while enhancing efficiency, simultaneously introduce challenges like hallucination, irreproducibility, data bias, and model opacity. To guide responsible adoption, the study proposes an interdisciplinary governance framework and a roadmap for explainable AI research in scholarly contexts.
Current literature indexing systems inadequately annotate publication types and study designs (PTs), hindering structured and comprehensive retrieval required for tasks such as evidence synthesis and systematic reviews. This work proposes a unified indexing framework specifically designed for PTs, moving beyond the conventional MeSH-based subject indexing paradigm. Centered on user objectives, the framework establishes a standardized classification schema and enables automated annotation of both full-text articles and bibliographic records—including preprints and unpublished manuscripts—through probabilistic scoring. Importantly, the approach is model-agnostic, emphasizing indexing logic and system architecture rather than reliance on specific machine learning models. It facilitates consistent cross-database indexing and supports automated query expansion, substantially improving retrieval efficiency in biomedical literature and significantly reducing the manual screening burden in evidence synthesis workflows.
This study addresses the challenge of organizing scientific knowledge amid the exponential growth of scholarly literature by proposing an automatic hierarchical classification method based on large language models (LLMs). Leveraging in-context learning (ICL) and prompt chaining, the approach performs three-level categorization—domain, discipline, and topic—within the Open Research Knowledge Graph (ORKG) taxonomy. The first systematic evaluation demonstrates that prompt chaining significantly outperforms conventional ICL, surpassing existing state-of-the-art models particularly at the domain and discipline levels. Although accuracy at the finest-grained topic level remains moderate (approximately 50%), this work validates that off-the-shelf LLMs, without fine-tuning, can effectively support hierarchical semantic organization of scientific texts through carefully engineered prompting strategies.