🤖 AI Summary
This study systematically compares the performance and applicability of encoder-only architectures (BERT) and sequence-to-sequence models (T5) on named entity recognition (NER) tasks. The authors evaluate BERT fine-tuned with weighted cross-entropy loss against T5 adapted via few-shot prompt-based learning, under both 7-class and 3-class label schemes. Through ablation studies, they analyze the impact of hyperparameters and characterize typical error patterns. The work presents a comprehensive, multi-metric comparison of the two modeling paradigms, elucidating their respective strengths and limitations. These findings offer empirical guidance and methodological insights for selecting and deploying NER models in practical applications.
📝 Abstract
Named entity recognition (NER) has been one of the essential preliminary steps in modern NLP applications. This report focuses on implementing the NER task on finetuning two pretrained models: (i) an encoder-only model (BERT) with a simple classification head, and (ii) a sequence-to-sequence model (T5) with few-shot prompts. Under the original 7-class tag and 3-class simplified tag schemes, BERT is applied a weighted cross-entropy for training loss, and T5 is fine-tuned with two validation strategies. It also conducted an ablation study with different hyperparameters. Moreover, the related analysis provides valuable insights into common errors in BERT and the two models' performance. Based on a bunch of performance metrics, this report aims to compare the above two architectures and explore their abilities in the sequence labelling task, laying the groundwork for further practical use cases.