🤖 AI Summary
This work addresses the challenge of nested named entity recognition (Nested NER), which typically relies on costly multi-layer annotations, while most existing corpora contain only flat annotations. The study presents the first systematic investigation into the feasibility of learning nested structures using solely flat annotations. It proposes a weakly supervised framework that integrates substring matching, pseudo-nested data generation, signal neutralization, and collaborative inference between fine-tuned models and large language models. Evaluated on the Russian NEREL dataset, the best-performing variant achieves an inner-span F1 score of 26.37%, closing 40% of the performance gap with fully supervised methods and substantially advancing the practicality of low-resource Nested NER.
📝 Abstract
Nested named entity recognition identifies entities contained within other entities, but requires expensive multi-level annotation. While flat NER corpora exist abundantly, nested resources remain scarce. We investigate whether models can learn nested structure from flat annotations alone, evaluating four approaches: string inclusions (substring matching), entity corruption (pseudo-nested data), flat neutralization (reducing false negative signal), and a hybrid fine-tuned + LLM pipeline. On NEREL, a Russian benchmark with 29 entity types where 21% of entities are nested, our best combined method achieves 26.37% inner F1, closing 40% of the gap to full nested supervision. Code is available at https://github.com/fulstock/Learning-from-Flat-Annotations.