scientific entity extraction

Designs and implements systems that identify and extract mentions of research-specific entities from scientific text, including span detection and named-entity recognition. Builds models and pipelines (e.g., NER or LLM-based extractors) to classify entity types, normalize variant mentions to canonical forms, and link mentions across documents to produce semantically coherent, scalable entity sets across diverse research areas.

scientificentityextraction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.48
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses a critical gap in academic information extraction by shifting focus from scholarly papers to implementation-level research artifacts in code repositories. We formally define and annotate ten categories of implementation-level entities within README files, introducing NERdME—a high-quality named entity recognition dataset comprising 200 expert-annotated READMEs and over 10,000 entity spans. Leveraging this resource, we conduct baseline experiments with large language models and fine-tuned Transformers, revealing substantial differences between implementation-level and paper-level entities. Furthermore, we demonstrate the practical utility of our approach through downstream entity linking tasks, showing its effectiveness in research artifact discovery and metadata integration. This study thus establishes the first systematic framework for semantic information extraction from code repositories, filling a longstanding void in academic knowledge extraction.

Code RepositoriesNamed Entity RecognitionREADME Files

LLM-Based Information Extraction to Support Scientific Literature Research and Publication Workflows

Oct 06, 2025
SA
Samy Ateia
🏛️ University of Regensburg | University of Bayreuth

With the exponential growth of scientific literature, automated extraction of key concepts remains challenging, particularly due to poor cross-disciplinary adaptability. Method: This paper proposes a lightweight LLM-based semantic extraction method supporting FAIR implementation in scholarly workflows. It introduces a context learning–driven zero-/few-shot domain adaptation mechanism that enables rapid, fine-tuning–free adaptation to new disciplines. We systematically benchmark multiple open-source and commercial LLMs on concept identification tasks and develop an interactive online prototype system. Contribution/Results: Empirical evaluation in computer science—complemented by user studies—demonstrates the method’s effectiveness in structured literature review, knowledge graph construction, and information retrieval. It significantly improves both accuracy and cross-domain generalization of concept extraction, offering a scalable technical pathway for intelligent, full-lifecycle scholarly knowledge services.

Enabling rapid domain adaptation for scientific information extractionExtracting key concepts from scientific documents using LLMsSupporting FAIR principles in scientific publishing workflows

Symbol-based entity marker highlighting for enhanced text mining in materials science with generative AI

May 09, 2025
JL
Junhyeong Lee
🏛️ Korea Institute of Energy Research | Hankyang University | Korea Advanced Institute of Science and Technology

To address the low accuracy of structured data extraction from materials science literature, this paper proposes a hybrid text-mining framework. First, symbolic entity markers are introduced to enhance named entity recognition (NER) performance; subsequently, a joint modeling approach integrates sequence labeling with structured generation to enable collaborative extraction of entities and relations. This method innovatively combines the strengths of multi-stage and end-to-end paradigms, overcoming traditional limitations in fine-grained entity identification and complex relational modeling. Evaluated on three authoritative benchmark datasets—MatScholar, SOFC, and one additional domain-specific corpus—the framework achieves a 58% improvement in entity-level F1 score and an 83% improvement in relation-level F1 score over state-of-the-art methods. The proposed approach establishes a new, efficient, and robust paradigm for constructing scientific literature knowledge graphs.

Enhanced entity recognition using symbolic annotationsHybrid text-mining framework for structured data conversionImproving entity and relation extraction in materials science

This work addresses the challenge of named entity recognition (NER) in scientific texts, where the large number of candidate entity types often hinders the performance of large language models. To mitigate this type overload issue, the authors propose TdSciNER, a type-driven multi-task learning framework that incorporates an auxiliary entity classification task. The approach further integrates a context example selection strategy based on sentence similarity and type diversity to enhance model generalization. Evaluated on three scientific text datasets, TdSciNER achieves performance comparable to fully supervised models, demonstrating that type filtering combined with the proposed example selection mechanism plays a crucial role in improving NER accuracy.

Entity Type ComplexityLarge Language ModelsPrompt-based Learning

Named Entity Analysis and Extraction with Uncommon Words

Oct 16, 2018
XZ
Xiaoshi Zhong
🏛️ Beijing Institute of Technology | Nanyang Technological University

This work addresses few-shot named entity recognition (NER) by challenging the end-to-end joint modeling paradigm. Drawing on generative grammar theory, we propose a “extract-then-classify” decoupled framework: entity extraction is treated as a syntactic task—requiring no semantic information—while classification is delegated to pre-trained language models (PLMs) or large language models (LLMs) as a semantic task. Empirical analysis reveals that rare words—particularly proper nouns—serve as critical syntactic cues; high-precision extraction is achieved using only shallow syntactic features (e.g., POS tags, dependency relations, and n-grams), with word embeddings or contextualized semantic representations yielding no performance gain. On benchmarks including CoNLL-2003, our extraction module achieves state-of-the-art F1 scores; ablation studies confirm that incorporating semantic features does not improve extraction accuracy. To our knowledge, this is the first study grounding syntactic–semantic separation in formal linguistics, providing both theoretical justification and empirical validation for decoupled modeling, while elucidating the root cause of failures in multi-task joint parsing.

Addressing few-shot named entity recognition with LLMsImproving NER performance using effective prompt constructionTransforming sequence-labeling into sequence-generation problem

Latest Papers

What's happening recently
View more

Inclusion of Role into Named Entity Recognition and Ranking

Nov 10, 2025
NK
Neelesh K. Shukla
🏛️ IIT Guwahati

This work addresses context-aware named entity role identification and ranking—i.e., identifying fine-grained semantic roles (e.g., “agent”, “beneficiary”) that entities assume within a given context, and subsequently retrieving and ranking semantically relevant entities. We model roles as domain-agnostic, fine-grained semantic subtypes and jointly learn role labels and entity representations via a lightweight sequence labeling framework. Furthermore, we introduce a role–entity semantic matching mechanism that integrates both sentence- and document-level contextual information for cross-granularity modeling. Our approach requires only a small amount of annotated data and achieves cross-domain generalization without domain-specific adaptation. Experiments demonstrate significant improvements over strong baselines in implicit role identification and low-resource settings. Moreover, the method exhibits strong robustness and transferability in role-driven entity retrieval tasks.

Detecting contextual roles of entities beyond basic named entity typesLearning role representations without large domain-specific datasetsRetrieving entity subsets based on their specific roles in text

GSAP-ERE: Fine-Grained Scholarly Entity and Relation Extraction Focused on Machine Learning

Nov 12, 2025
WO
Wolfgang Otto
🏛️ GESIS – Leibniz Institute for the Social Sciences

To address the challenges of fine-grained concept and relation extraction from machine learning (ML) scholarly literature—and weak support for reproducibility—this paper introduces ML-AIE, the first high-quality academic information extraction dataset specifically designed for ML. It covers the full text of 100 peer-reviewed papers and contains 63,000 manually annotated entities (10 types) and 35,000 semantic relations (18 types). Leveraging this human-annotated corpus, fine-tuned models achieve 80.6% and 54.0% F1 scores on named entity recognition (NER) and relation extraction (RE), respectively—substantially outperforming state-of-the-art large language model prompting approaches (44.4% and 10.1%). ML-AIE fills a critical gap in benchmarking fine-grained academic knowledge extraction for ML, serving as foundational infrastructure and an evaluation standard for constructing ML knowledge graphs, monitoring research reproducibility, and developing domain-specific information extraction models.

Building dataset for knowledge graph construction and reproducibility monitoringEvaluating performance gaps between fine-tuned models and LLM methodsExtracting fine-grained entities and relations from ML publications

Structure-Aware Decoding Mechanisms for Complex Entity Extraction with Large-Scale Language Models

Dec 15, 2025
ZQ
Zhimin Qiu
🏛️ University of Southern California | Washington University in St. Louis | Stevens Institute of Technology | University of Pennsylvania

To address the challenge of jointly preserving semantic integrity and structural consistency in nested and overlapping named entity recognition (NER), this paper proposes a structure-aware decoding framework. The method leverages pretrained language model representations and integrates multi-granularity span composition with hierarchical decoding. Its core contributions are: (1) a novel collaborative mechanism between candidate span generation and structured attention, explicitly modeling entity boundaries, hierarchical nesting, and cross-entity dependencies; and (2) a joint optimization objective incorporating hierarchical structural constraints and semantic–structural consistency, combining classification loss with structure-consistency loss. Experimental results on ACE 2005 demonstrate significant F1-score improvements over prior work. The approach achieves superior precision, recall, and boundary localization for both nested and overlapping entities, and exhibits strong robustness on long sentences and scenarios with multiple co-occurring entities.

Extracts nested and overlapping entities with structural consistencyMaintains semantic integrity in complex entity recognition tasksModels entity boundaries, hierarchies, and cross-dependencies simultaneously

This work addresses the absence of high-quality, multi-domain named entity recognition and linking (NERL) datasets for historical Italian by introducing ENEIDE, the first publicly available silver-standard dataset for this language variety. ENEIDE comprises 2,111 documents from two scholarly digital collections spanning the 18th to 20th centuries, annotated with over 8,000 entities and partitioned into training, development, and test sets. The dataset incorporates an innovative NIL (not-in-lexicon) handling mechanism and leverages semi-automatic annotation, Wikidata entity linking, and rigorous quality control. Baseline experiments demonstrate that ENEIDE poses a substantial challenge to current NERL models, revealing a significant performance gap between zero-shot and fine-tuned approaches, while also enabling temporal disambiguation and cross-domain evaluation.

Entity DisambiguationHistorical ItalianNamed Entity Linking

This work addresses the scarcity of fine-grained annotated resources for named entity recognition (NER) and relation extraction (RE) in art history. To bridge this gap, the authors introduce FRAME, a novel dataset comprising descriptions of individual artworks sourced from museum catalogs and auction records. The dataset features a three-tier manual annotation scheme—metadata, content, and coreference layers—covering 37 entity types aligned with Wikidata. Annotations are provided in stand-off format using UIMA XMI CAS, facilitating tasks such as entity linking, knowledge graph construction, and fine-tuning or evaluation of large language models. As the first open-source, fine-grained, multi-layer annotated resource tailored to art historical research, FRAME establishes a benchmark for NER, RE, and few-shot or zero-shot modeling in this domain.

Art-historical Image DescriptionsDatasetEntity Annotation

Hot Scholars

CZ

Chengzhi Zhang

Nanjing University of Science and Technology
Text MiningNatural Language ProcessingScience of Science
DH

Dominik Heider

Director, University of Münster
Data ScienceMachine LearningArtificial IntelligenceBiomedical Informatics
GH

Georges Hattab

Adjunct Professor of Computer Science, Freie Universität Berlin, Robert Koch Institute
Artificial IntelligenceData MiningVisualization
RK

Rohan Kumar

Carnegie Mellon University
Natural Language ProcessingDeep LearningInformation Retrieval
XZ

Xuanhe Zhou

Assistant Professor, Shanghai Jiao Tong University
Data ManagementArtificial Intelligence