Consensus Measures for Unstructured Biomedical Text Annotations

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of quantifying soft inter-annotator agreement in open-label biomedical text annotation, where traditional metrics fall short due to the semantic flexibility of labels. To tackle this issue, the authors propose a semantic equivalence measure grounded in natural language inference (NLI), which effectively balances scalability with fine-grained conceptual discrimination. By integrating word embeddings, large language models, and NLI techniques, the method captures semantic similarity among open-ended labels in unstructured text. Synthetic experiments demonstrate that multiple semantic metrics can quantify soft agreement, yet they exhibit substantial differences in estimation bias. These findings underscore the novelty and practical utility of the proposed approach for evaluating annotation quality in complex, open-label settings.
📝 Abstract
Biomedical literature is increasingly mined for knowledge beyond the questions it was written to answer. Because the target concepts are not known in advance, annotators prefer open-ended labels, whose agreement is hard to quantify. We study soft inter-rater reliability for annotators providing unstructured texts for biomedical annotation tasks. Synthetic experiments show that soft reliability can be quantified using a variety of semantic equivalence measures, and that the choice of measure affects failure modes of the estimation. Embeddings are scalable, but limited when differentiating similar but distinct concepts. Large language models are promising, but limited by scalability for estimating agreement by chance. Finally, we suggest measures based on natural language inference as a sensible compromise.
Problem

Research questions and friction points this paper is trying to address.

inter-rater reliability
unstructured text
biomedical annotation
semantic equivalence
open-ended labels
Innovation

Methods, ideas, or system contributions that make the work stand out.

soft inter-rater reliability
semantic equivalence
natural language inference
unstructured biomedical annotation
large language models
🔎 Similar Papers
No similar papers found.