🤖 AI Summary
This study addresses the limited robustness in biomedical named entity recognition caused by weak structural constraints in instruction serialization and the scarcity of high-quality annotations. To tackle these challenges, this work proposes MITE, a method that reformulates the task as code-formatted structure-to-structure generation. By leveraging representations from multiple programming languages—including Python, C++, and Java—MITE provides diverse structured supervision signals. Furthermore, it introduces an entity-level voting mechanism to aggregate predictions, operating entirely without external knowledge or additional annotations. Experimental results demonstrate that MITE significantly outperforms both BERT-based and large language model baselines across six benchmark datasets, exhibiting superior cross-domain generalization capability and robustness.
📝 Abstract
Instruction tuning has become a common paradigm for applying large language models (LLMs) to biomedical named entity recognition (BioNER). However, existing instruction-tuning approaches still face two key challenges. First, conventional natural-language instructions typically serialize BioNER annotations as flat textual outputs, providing limited structural constraints for typed entity extraction. Second, high-quality biomedical annotations are limited, and learning from a single serialized output form may restrict structural diversity and reduce model robustness. Although external biomedical knowledge can be introduced to alleviate data scarcity, it often requires costly resource construction. To address these challenges, we propose MITE, a Multiple Programming Languages Instruction Tuning and Ensemble method for BioNER. MITE reformulates BioNER as a structure-to-structure generation task by representing both instructions and entity outputs in code-formatted representations. Specifically, each training instance is transformed into multiple programming-language formats, including Python, C++, and Java, while preserving the same underlying entity semantics. These language-specific representations provide structurally diverse supervision without requiring external biomedical knowledge or additional annotations. During inference, MITE aggregates predictions from different code formats through an entity-level voting strategy, reducing language-specific prediction variance and improving robustness. Experiments on six widely used BioNER datasets demonstrate that MITE consistently outperforms representative BERT-based and LLM-based baselines and exhibits strong cross-dataset generalization. Ablation and parameter analyses further verify the effectiveness and robustness of the proposed components.