A Guideline-Augmented Multi-Agent Framework for Schema-as-Code Biomedical Named Entity Recognition

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of annotation guidelines and uncontrollable generation structures when applying large language models (LLMs) to biomedical named entity recognition (BioNER). To tackle these challenges, we propose GAMA, a multi-agent framework that introduces a novel Schema-as-Code paradigm. Specifically, GAMA constructs a guideline memory by inducing and validating annotation rules from training data, integrating planning, coding, and verification modules to achieve structured extraction. Furthermore, a dual-loop refinement mechanism is designed to effectively correct model hallucinations and boundary errors. Experimental results demonstrate that our approach significantly outperforms existing baselines across five mainstream BioNER datasets. Ablation studies further confirm the effectiveness of each component in enhancing both prediction accuracy and format compliance.
📝 Abstract
Large language models (LLMs) have shown promising potential for biomedical named entity recognition (BioNER) through instruction following and in-context learning. However, existing LLM-based BioNER methods still face two key limitations. First, retrieved demonstrations and external biomedical knowledge provide limited support for dataset-specific annotation semantics, leaving entity boundaries, type scopes, and annotation conventions ambiguous. Second, free-form generation lacks sufficient structural control, often leading to invalid formats, hallucinated mentions, duplicated entities, and boundary errors. To address these limitations, we propose GAMA, a guideline-augmented multi-agent framework for schema-as-code BioNER. GAMA first induces candidate annotation rules from labeled training instances and verifies them against annotated data to construct reliable dataset-specific guideline memory. Guided by these verified rules, a planning component generates ranked span-type hypotheses with rationales, and a coding component converts them into schema-constrained entity objects. A verification module then checks span grounding, type validity, and structural compliance, and performs dual-loop refinement to correct invalid or low-confidence predictions. Experiments on five widely used BioNER datasets with multiple LLM backbones show that GAMA consistently outperforms strong LLM-based baselines. Ablation and parameter analyses further verify the effectiveness of the proposed components.
Problem

Research questions and friction points this paper is trying to address.

Biomedical Named Entity Recognition
Large Language Models
Annotation Guidelines
Structural Control
Hallucination
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Framework
Schema-as-Code
Guideline-Augmented
Biomedical Named Entity Recognition
Dual-loop Refinement