Score
Designs and implements systems and analytic workflows to collect, aggregate, and analyze scholarly outputs and related signals (e.g., publications, preprints, patents, grants, citations) to detect and characterize emerging topics, shifts, and bursts in research activity. Builds taxonomies, metrics, automated alerts, dashboards, and statistical or machine‑learning analyses (e.g., topic modeling, burst detection, trend forecasting) to monitor, visualize, and report research trends over time.
Existing approaches extract program components in isolation, failing to reconstruct complete scientific workflows—thereby impeding research reproducibility and the advancement of “AI for Science.” This paper introduces the first end-to-end, paper-level workflow generation framework that integrates paragraph-level text mining with generative modeling to automatically construct structured, source-locatable, and visualizable research flowcharts from full-text papers. Methodologically, it employs SciBERT coupled with PU learning to identify descriptive paragraphs; leverages Flan-T5 with prompt engineering to generate workflow phrases; and applies few-shot learning via ChatGPT for stage classification and precise mapping to original text locations. Evaluated on NLP-domain papers, the framework achieves a paragraph identification F1-score of 0.977, ROUGE-1 of 0.454 for workflow phrase generation, and 95.8% accuracy in stage classification. It further enables the first systematic, longitudinal analysis of methodological evolution in NLP over the past two decades—revealing marked growth in data analysis and ablation studies.
Existing bibliometric tools exhibit high rigidity, limited flexibility, and steep technical barriers, requiring researchers to possess programming expertise for dynamic scientometric analysis. To address this, we propose the first generative multi-agent AI framework specifically designed for scientometrics, enabling end-to-end analysis via natural language instructions—without coding prerequisites. Our framework innovatively integrates natural language–to–code translation, multimodal full-text retrieval, autonomous agent exploration, and dynamic metric construction. It employs a four-agent collaborative architecture: a Custom Analysis Generator, a Full-Text Retriever, a RAG-powered Research Assistant, and an Automated Report Generator—augmented with sandboxed execution, topic modeling, and embedding-based clustering. Experimental evaluation demonstrates that the system autonomously generates executable scripts, accurately identifies research frontiers, constructs collaboration and citation networks, and produces reproducible, structured scientific reports.
This study addresses the challenges of dynamic academic trajectory analysis and weak causal attribution of career evolution among researchers. Methodologically, we propose an interactive visual analytics system for academic career modeling, built upon a multidimensional AI framework: (1) integrating BERTopic/LDA with temporal UMAP for topic evolution modeling; (2) employing graph neural networks to capture dynamic co-authorship network evolution; and (3) introducing a novel configurable prompt-driven large language model (LLM) mechanism to automatically generate structured, causal narratives of career progression. The system enables cross-scholar comparison, milestone-oriented causal tracing, and millisecond-level coordinated interactions across multiple coordinated views. Evaluated on real-world scholar datasets, our approach achieves 92% accuracy in topic transition detection, and the generated reports are validated by domain experts. This work delivers an interpretable, interactive, and intelligent analytical infrastructure for scholarly assessment and career planning.
The exponential growth of academic literature has severely hampered manual scholarly discovery. To address this, we propose Agent-E—the first end-to-end system integrating task-oriented AI agents with robotic process automation (RPA) to automatically identify geographically relevant research findings from conference proceedings and trigger downstream actions (e.g., award nominations). Methodologically, Agent-E combines named entity recognition, fine-grained geographic coding, and RPA to precisely extract and act upon geographic intelligence. Evaluated on 586 papers across five major conferences, it achieves 100% recall and 99.4% precision for target papers. This work pioneers the deep synergistic integration of AI agents and RPA for automated academic geointelligence, significantly enhancing research administration efficiency. It establishes a reusable technical paradigm for intelligent, domain-aware academic workflows.
Amid accelerating scientific progress and information overload, evaluating research impact and allocating resources effectively remains challenging. Method: We propose and validate a novel predictive framework for identifying highly cited papers early, leveraging interdisciplinary classification models trained on heterogeneous academic data from computer science, physics, and PubMed. The framework extracts paper-level features—including citation dynamics, thematic evolution, and author collaboration networks—to forecast high-impact publications within the 2010–2024 period. Contribution/Results: Experiments demonstrate consistently strong predictive performance across domains (AUC > 0.85), providing the first systematic empirical evidence that high-impact scientific work exhibits observable, quantifiable structural precursors. By uncovering universal statistical regularities underlying scientific discovery, our framework delivers an interpretable, reusable, and domain-agnostic quantitative foundation for evidence-based research policy, peer review, and funding allocation.
This work addresses the lack of effective monitoring of dataset usage in scholarly literature, which undermines citation transparency, impact traceability, and reproducibility. To tackle this challenge, the study introduces the first application of the multi-task GLiNER framework to dataset usage monitoring, jointly performing dataset mention extraction, relation identification, and usage context classification. The approach integrates synthetic data generation with a large language model (LLM)-driven re-verification mechanism to mitigate issues of annotation scarcity and ambiguous citations. This combination significantly enhances the accuracy, coverage, and label consistency of dataset mention detection, enabling end-to-end, unconstrained tracking of data citations across diverse scientific texts and advancing the development of open-source tools for scholarly data provenance.
This study investigates how academic age influences methodological choices among scholars in Library and Information Science (LIS). Drawing on a corpus of 26,677 articles published between 1990 and 2023 in 14 leading journals, the authors employ author disambiguation techniques to compute academic age and introduce, for the first time, the CogFT model to automatically classify research methods, complemented by Top2Vec for topic modeling. The analysis reveals a sustained decline in the use of theoretical approaches alongside significant increases in experimental and bibliometric methods. Methodological diversity peaks among mid-career researchers and is lowest among late-career scholars. These findings illuminate evolving patterns in research methodology within LIS and offer empirical guidance for early-career researchers in selecting appropriate methods.
This work addresses critical limitations of traditional science, technology, and innovation (STI) analysis—namely its reliance on static indicators that suffer from time lags, shallow semantic representation, and an inability to capture the nonlinear dynamics of knowledge ecosystems. To overcome these challenges, the study proposes a validation-centric, five-layer hybrid framework that constructs a dynamic, versioned knowledge graph from open scholarly data and integrates constrained large language models (LLMs) for structured semantic enrichment. Reliability is ensured through a multi-tiered validation pipeline combining structural, evidential, comparative, and expert-based verification, augmented by a traceability mechanism. This approach significantly enhances the semantic depth, timeliness, and credibility of STI analysis while upholding scientific evidentiary standards, thereby enabling robust detection of emerging trends, mapping of technology transfer pathways, and policy-relevant gap analysis.
Existing datasets of scientific ideation trajectories struggle to comprehensively capture the full research process—from literature exploration and tool utilization to the evolution of intermediate artifacts and final proposals. This work proposes a reverse-to-forward synthesis mechanism that emulates the uncertainty, evidence integration, and phased convergence characteristic of real scientific inquiry through a Generator–Advisor architecture. By leveraging action–observation–editing sequence modeling, context-aware verification, and process-level supervision, the approach generates multi-turn trajectories aligned with authentic research practices, starting from high-quality papers and proposals. The study yields the first trajectory dataset spanning the complete scientific workflow and establishes a generalizable paradigm for synthesizing process-supervised data for scientific agents.