Named Entity Recognition in COVID-19 tweets with Entity Knowledge Augmentation

📅 2025-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address challenges in named entity recognition (NER) for COVID-19 social media text—including informal language, scarce annotated data, and strong domain knowledge dependency—this paper proposes an entity knowledge-enhanced framework. Built upon a pre-trained language model, the framework integrates external biomedical knowledge (e.g., UMLS ontology) and entity prior information into input representations via lightweight knowledge injection, enabling end-to-end training under both few-shot and fully supervised settings. Experiments on COVID-19 Twitter and PubMed datasets demonstrate substantial performance gains: F1 score improves by +8.2% in few-shot scenarios. Moreover, the method exhibits strong transferability to general biomedical NER tasks. The core contribution lies in a lightweight, scalable knowledge fusion mechanism that jointly ensures robustness and generalization without architectural complexity.

Technology Category

Natural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)Data Mining & Knowledge Management: Linked Open Data, Knowledge Graphs & KB CompletionKnowledge Representation and Reasoning: Knowledge Engineering

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
The COVID-19 pandemic causes severe social and economic disruption around the world, raising various subjects that are discussed over social media. Identifying pandemic-related named entities as expressed on social media is fundamental and important to understand the discussions about the pandemic. However, there is limited work on named entity recognition on this topic due to the following challenges: 1) COVID-19 texts in social media are informal and their annotations are rare and insufficient to train a robust recognition model, and 2) named entity recognition in COVID-19 requires extensive domain-specific knowledge. To address these issues, we propose a novel entity knowledge augmentation approach for COVID-19, which can also be applied in general biomedical named entity recognition in both informal text format and formal text format. Experiments carried out on the COVID-19 tweets dataset and PubMed dataset show that our proposed entity knowledge augmentation improves NER performance in both fully-supervised and few-shot settings. Our source code is publicly available: https://github.com/kkkenshi/LLM-EKA/tree/master
Problem

Research questions and friction points this paper is trying to address.

Recognizing named entities in informal COVID-19 tweets
Addressing limited annotated data for robust NER models
Incorporating domain-specific knowledge for biomedical NER
Innovation

Methods, ideas, or system contributions that make the work stand out.

Entity knowledge augmentation for COVID-19 NER
Improves NER in supervised and few-shot settings
Applicable to informal and formal biomedical texts
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xuankang Zhang
School of Information Science and Engineering, Yunnan University, China
Jiangming Liu
Jiangming Liu
Associate Professor, Yunnan University
natural language processingdeep learning