literature mining

Designs and implements pipelines and tools to extract, normalize, and curate structured information from corpora of scholarly texts—for example citation relation types, argumentative roles, and domain-specific analytical priors. Builds datasets and processing components that link, filter, and prepare literature-derived facts and metadata for downstream analysis, search, or modeling.

literaturemining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.08
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a novel dataset discovery framework that leverages citation contexts from scientific papers to better capture the semantic intent behind research queries, addressing the limitations of existing dataset search engines that rely primarily on metadata and keyword matching and consequently suffer from low recall. By treating citation context as the core signal—combined with large-scale context extraction, large language model–guided pattern recognition, and provenance-preserving entity resolution—the approach significantly reduces dependence on incomplete or inconsistent metadata. Evaluated on eight computer science queries, the method achieves an average normalized recall of 47.47% (peaking at 81.82%), substantially outperforming Google Dataset Search and DataCite Commons. The framework’s novelty and practical utility have been affirmed by domain experts across multiple disciplines.

citation contextdataset discoverymetadata

Facets, Taxonomies, and Syntheses: Navigating Structured Representations in LLM-Assisted Literature Review

Apr 25, 2025
RF
Raymond Fok
🏛️ University of Washington | Allen Institute for AI

Large-scale literature reviews face significant challenges in automated deep analysis and synthesis due to insufficient semantic understanding and structural reasoning capabilities. To address this, we propose DimInd, an interactive system introducing a novel hierarchical compression-based structured representation framework. It unifies paper-level understanding, multi-dimensional comparison, conceptual categorization, and narrative synthesis into a traceable, progressive workflow: papers → comparative tables → conceptual taxonomy → narrative review. DimInd integrates prompt engineering, structured information extraction, hierarchical clustering modeling, and interactive visualization, leveraging large language models (LLMs) for end-to-end semantic parsing and organization. In evaluations with 23 researchers, DimInd significantly reduced cognitive load in information extraction and conceptual organization compared to a ChatGPT baseline, while improving review construction efficiency and structural coherence. It is the first system to enable automated, deep, and narratively coherent synthesis for large-scale scholarly corpora.

Facilitates synthesis of large paper collections for literature reviewsProvides structured representations to guide literature understandingReduces manual effort in organizing and analyzing research papers

The Semantic Scholar Open Data Platform

Apr 11, 2026
RM
Rodney Michael Kinney
🏛️ Allen Institute for Artificial Intelligence

Amidst the exponential growth of scientific literature, researchers urgently require efficient tools for literature understanding and discovery. This paper introduces Semantic Scholar’s open academic knowledge graph construction paradigm: a novel, fully automated pipeline integrating multi-source data, high-precision PDF parsing, fine-grained structured semantic annotation, NLP-driven natural language summarization, and context-aware embedding representation learning. The resulting open academic graph—the largest to date—comprises over 200 million papers, 80 million authors, and 2.4 billion citations, hosted on a dynamically updatable “living document”–style platform architecture. We publicly release the Semantic Scholar Academic Graph alongside standardized APIs, establishing it as a globally adopted open research infrastructure. This framework significantly enhances the efficiency and effectiveness of scholarly information retrieval, comprehension, and knowledge synthesis.

Automated tools needed to manage growing scientific literature volumeBuilding large open academic graph with advanced semantic featuresSemantic Scholar accelerates science by enhancing literature discovery

SciDaSynth: Interactive Structured Knowledge Extraction and Synthesis from Scientific Literature with Large Language Model

Apr 21, 2024
XW
Xingbo Wang
🏛️ Weill Cornell Medicine | Cornell University | Hong Kong University of Science and Technology

Scientific literature is inherently multimodal, heterogeneous, and unstructured, posing significant challenges for existing knowledge extraction systems in achieving cross-document consistency and dynamic adaptation to user intent. To address this, we propose the first LLM-driven interactive knowledge structuring paradigm, integrating prompt engineering, structured output control, conversational state management, and multi-granularity visual exploration. This enables researchers to automatically generate structured tables via natural language queries while collaboratively verifying and iteratively refining outputs. Our approach overcomes key limitations of conventional automated systems: it maintains high accuracy and coverage while reducing manual correction effort by over 40%. Empirical evaluation demonstrates substantial improvements in the efficiency of constructing high-quality scientific knowledge bases, offering a novel paradigm for domain-specific knowledge graph construction and reproducible research.

Building scalable interactive systems for literature-based synthesisExtracting structured knowledge from multimodal scientific literatureProcessing inconsistent information across diverse research papers

SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding

Aug 28, 2024
SL
Sihang Li
🏛️ University of Science and Technology of China | DP Technology

Large language models (LLMs) face dual challenges in scientific literature understanding: insufficient domain-specific knowledge and poor task alignment. To address these, we propose a “knowledge injection–task alignment” collaborative adaptation framework. Our method introduces a novel scientific text quality enhancement pipeline and constructs SciLitIns—the first high-quality instruction dataset tailored to niche scientific domains—generated via an LLM-driven synthetic instruction approach. The pipeline integrates robust PDF parsing, multi-stage quality filtering, continued pretraining (CPT), and supervised fine-tuning (SFT). The resulting model, SciLitLLM, achieves significant performance gains over general-purpose baselines across multiple scientific literature understanding benchmarks, empirically validating the efficacy of synergistic knowledge enhancement and task-specific refinement. Moreover, the framework demonstrates cross-domain transferability, establishing a systematic paradigm for adapting foundation models to specialized scientific domains.

Adapt LLMs for scientific literature understandingEnhance LLMs' scientific domain knowledgeImprove LLMs' performance on specialized scientific tasks

Latest Papers

What's happening recently
View more

This work addresses the lack of effective monitoring of dataset usage in scholarly literature, which undermines citation transparency, impact traceability, and reproducibility. To tackle this challenge, the study introduces the first application of the multi-task GLiNER framework to dataset usage monitoring, jointly performing dataset mention extraction, relation identification, and usage context classification. The approach integrates synthetic data generation with a large language model (LLM)-driven re-verification mechanism to mitigate issues of annotation scarcity and ambiguous citations. This combination significantly enhances the accuracy, coverage, and label consistency of dataset mention detection, enabling end-to-end, unconstrained tracking of data citations across diverse scientific texts and advancing the development of open-source tools for scholarly data provenance.

academic data trackingdata citationdataset usage monitoring

Existing academic data systems struggle to uniformly support diverse query types—such as retrieval, knowledge discovery, and generation—and lack interpretable execution mechanisms. This work proposes an intelligent data management system tailored for academic corpora, which automatically compiles natural language queries into interpretable directed acyclic graph (DAG) execution plans. The system integrates structure-aware knowledge representation, large language model–driven hybrid query planning, and a unified execution framework based on composable operators. By synergistically combining structured knowledge management, agent-based planning, and explainable execution, the approach supports the full spectrum of academic queries and significantly outperforms existing systems in effectiveness, efficiency, and interpretability, thereby establishing a practical foundation for agent-driven academic data management.

data managementknowledge representationquery processing

Exploring LLMs for Scientific Information Extraction Using The SciEx Framework

Dec 10, 2025
SL
Sha Li
🏛️ Virginia Tech | University of Michigan

To address three core challenges in scientific literature information extraction—modeling long documents, understanding multimodal content, and standardizing fine-grained cross-paper information (especially under dynamically evolving data schemas)—this paper proposes SciEx, a modular, decoupled framework. SciEx explicitly separates PDF parsing, multimodal retrieval, LLM-driven extraction, and cross-document aggregation, enabling plug-and-play integration of diverse prompting strategies, foundation models, and inference mechanisms for rapid adaptation. Evaluated across three domain-specific datasets, SciEx achieves high accuracy and consistency in fine-grained information extraction. The study systematically identifies key strengths and bottlenecks of current LLM-based pipelines, offering an extensible and maintainable technical pathway for constructing scientific knowledge graphs that evolve with shifting data patterns and scholarly conventions.

Adapting extraction systems to rapidly changing data schemas or ontologies.Extracting fine-grained scientific data from long, multi-modal documents.Reconciling inconsistent information across publications into standardized formats.

This work addresses the challenge of efficiently accessing structured scholarly publications and associated software metadata within research knowledge bases. To this end, the authors propose a generic and extensible interoperable pipeline architecture built upon the shared Grid’5000/ABACA infrastructure, integrating modules for document preprocessing, information extraction, software mention recognition, and visualization. Designed to support multi-team collaboration, user validation, and external interoperability, the system demonstrates its utility through a daily tracking application of software mentions in the HAL open archive, significantly enhancing the visibility of research software and advancing open science practices. Experimental results confirm that the pipeline efficiently processes large-scale scientific literature and enables automated extraction and visual representation of software mentions within the HAL portal.

open scienceresearch repositoriesscientific literature

Hot Scholars

CZ

Chengzhi Zhang

Nanjing University of Science and Technology
Text MiningNatural Language ProcessingScience of Science
YL

Yuyu Luo

Assistant Professor, HKUST(GZ) / HKUST
Data AgentsLLM AgentsDatabaseText-to-SQL
HX

Hongyi Xu

Associate Professor at University of Connecticut | Ford R&A | '14 PhD, Northwestern
Engineering DesignDigital ManufacturingArtificial IntelligenceMicrostructure
CZ

Chuxu Zhang

Associate Professor of CSE, University of Connecticut (UConn)
Machine LearningDeep LearningData Mining