INSIGHTBUDDY-AI: Medication Extraction and Entity Linking using Large Language Models and Ensemble Learning

๐Ÿ“… 2024-09-28
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 1
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenges of extracting and standardizing drug-related information (e.g., dosage, route of administration, strength, adverse reactions) from clinical text. To this end, we propose the first multi-LLM stacking/voting ensemble framework for end-to-end drug entity recognition, normalization, and cross-knowledge-base mappingโ€”supporting SNOMED-CT, BNF, dm+d, and ICD. Our method integrates domain-finetuned BERT models (BioClinicalBERT, PubMedBERT), rule-augmented entity linking, and collaborative reasoning from multiple LLMs (LLaMA, ChatGLM). Evaluated on both general and clinical-specific benchmarks, our approach achieves superior F1-scores over state-of-the-art medical NER models. We release an open-source toolkit enabling lightweight deployment and have successfully integrated it into real-world hospital NLP workflows, significantly enhancing the structural fidelity and semantic interoperability of drug information extraction.

Technology Category

Natural Language Processing: Information ExtractionMachine Learning: Large Multimodal Models (LMMs)Data Mining & Knowledge Management: Linked Open Data, Knowledge Graphs & KB Completion

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
๐Ÿ“ Abstract
Medication Extraction and Mining play an important role in healthcare NLP research due to its practical applications in hospital settings, such as their mapping into standard clinical knowledge bases (SNOMED-CT, BNF, etc.). In this work, we investigate state-of-the-art LLMs in text mining tasks on medications and their related attributes such as dosage, route, strength, and adverse effects. In addition, we explore different ensemble learning methods ( extsc{Stack-Ensemble} and extsc{Voting-Ensemble}) to augment the model performances from individual LLMs. Our ensemble learning result demonstrated better performances than individually fine-tuned base models BERT, RoBERTa, RoBERTa-L, BioBERT, BioClinicalBERT, BioMedRoBERTa, ClinicalBERT, and PubMedBERT across general and specific domains. Finally, we build up an entity linking function to map extracted medical terminologies into the SNOMED-CT codes and the British National Formulary (BNF) codes, which are further mapped to the Dictionary of Medicines and Devices (dm+d), and ICD. Our model's toolkit and desktop applications are publicly available (at url{https://github.com/HECTA-UoM/ensemble-NER}).
Problem

Research questions and friction points this paper is trying to address.

Drug Information Extraction
Text Analysis
Medical Database Integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Supercomputing
Ensemble Models
Automated Coding in Healthcare
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Manchester Metropolitan University | Leiden University | University of Manchester