fine-tune ner models

Designs and implements training pipelines that adapt pretrained sequence-labeling models for named entity recognition on a target corpus, including dataset preparation, label mapping, annotation formats, transfer-learning or domain-adaptation strategies, and hyperparameter tuning. Builds and analyzes fine-tuned NER models by running controlled experiments and evaluating performance with sequence-labeling metrics (e.g., precision, recall, F1) to quantify expected gains and failure modes.

fine-tunenermodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.56
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of adapting traditional sequence-labeling-based named entity recognition (NER) to the generative paradigm of large language models (LLMs), and presents the first systematic evaluation of open-source LLMs on both flat and nested NER tasks. Through experiments on standard benchmarks using parameter-efficient fine-tuning (e.g., LoRA), structured output formats (inline brackets, XML), and models across multiple scales, the work demonstrates that open-source LLMs can achieve performance comparable to—or even surpassing—that of conventional encoder-based models and GPT-3 when prompted with structured formats. The results indicate that the NER capability of these models stems from their instruction-following and generalization abilities rather than memorization of entity–label pairs, and that fine-tuning has minimal adverse impact on their general capabilities—sometimes even enhancing them.

Fine-tuningGenerative NERLarge Language Models

This study systematically compares the performance and applicability of encoder-only architectures (BERT) and sequence-to-sequence models (T5) on named entity recognition (NER) tasks. The authors evaluate BERT fine-tuned with weighted cross-entropy loss against T5 adapted via few-shot prompt-based learning, under both 7-class and 3-class label schemes. Through ablation studies, they analyze the impact of hyperparameters and characterize typical error patterns. The work presents a comprehensive, multi-metric comparison of the two modeling paradigms, elucidating their respective strengths and limitations. These findings offer empirical guidance and methodological insights for selecting and deploying NER models in practical applications.

BERTModel ComparisonNamed Entity Recognition

Nested Named Entity Recognition as Single-Pass Sequence Labeling

May 22, 2025
AM
Alberto Munoz-Ortiz
🏛️ Universidade da Coruña | INSA Rennes | IRISA | Inria | CNRS | Université de Rennes

This work addresses the challenges of high computational complexity and structural intricacy in nested named entity recognition (NNER). We propose a single-pass sequence labeling approach that linearizes nested entity structures—represented as constituent trees—into token-level label sequences, thereby achieving the first complete reduction of NNER to standard token classification. Unlike prior methods, our approach eliminates the need for span enumeration, hierarchical decoding, or graph-based modeling; instead, it relies solely on a pretrained encoder (e.g., BERT) and a lightweight linearization strategy, ensuring seamless integration with mainstream sequence labeling frameworks. Crucially, it preserves expressive power while reducing inference time complexity to *O(n)*, significantly improving training and deployment efficiency. Extensive experiments demonstrate state-of-the-art performance across multiple benchmarks, validating the method’s effectiveness, simplicity, and strong generalization capability.

Achieves competitive performance efficientlyReduces nested NER to sequence labelingUses constituency linearizations with pretrained encoders

Open Named Entity Recognition (Open NER) suffers from weak cross-dataset and cross-lingual generalization, compounded by inconsistent entity category definitions across resources. To address these challenges, we propose B2NER: a unified framework comprising three core components. First, we construct the first comprehensive, fine-grained entity taxonomy covering 400+ categories. Second, we introduce a two-stage data refinement paradigm—standardizing heterogeneous category schemas via taxonomy alignment, followed by diversity-driven data pruning based on semantic clustering and coverage optimization. Third, we employ lightweight supervised fine-tuning to adapt large language models efficiently. B2NER achieves taxonomy-consistent modeling across 54 Chinese and English NER datasets. Experiments demonstrate that B2NER consistently outperforms GPT-4 by +6.8–12.0 F1 points across 15 datasets and six languages on three cross-domain benchmarks, significantly advancing zero-shot and few-shot Open NER performance.

Entity Category ConsistencyGeneralization AbilityOpen Named Entity Recognition

Latest Papers

What's happening recently
View more

This work addresses the challenges of multi-domain named entity recognition (NER) under conditions of scarce labeled data, where conventional approaches suffer from domain discrepancy, data sparsity, and overfitting. To overcome these limitations, the paper proposes a unified framework that integrates unsupervised pre-training, transfer learning, data augmentation, few-shot learning, and domain-adversarial training. This approach enables effective adaptation to target domains without requiring any annotated data therein, significantly enhancing model generalization and robustness in low-resource, multi-domain settings. Experimental results demonstrate that the proposed method substantially outperforms existing baselines, offering a novel and practical pathway toward efficient and transferable NER in resource-constrained environments.

Data ScarcityMultidomainNamed Entity Recognition

This study addresses the persistent performance gap between large language models (LLMs) and supervised fine-tuned models in named entity recognition (NER), noting that few-shot in-context learning fails to harness the full potential of abundant examples. The authors systematically investigate many-shot in-context learning for NER and demonstrate, for the first time, that providing hundreds of contextual examples enables LLMs to match or even surpass the performance of fine-tuned BERT models. Building on this finding, they propose a novel high-quality data annotation paradigm tailored for low-resource settings: by augmenting only around one hundred manually labeled samples to generate enriched training corpora, the F1 score of fine-tuned BERT improves by approximately 10%.

In-Context LearningLarge Language ModelsLow-Resource Annotation

This work addresses the lack of systematic evaluation of key design choices in multilingual named entity recognition (NER) models, which has obscured the true contributions of individual components to overall performance. The study presents the first comprehensive disentangled analysis across multiple dimensions—including model architecture, multilingual Transformer backbones, training objectives, and data composition—supported by large-scale ablation experiments. Based on these insights, the authors develop Otter, an efficient and general-purpose NER model supporting over 100 languages. Otter achieves a 5.3-point F1 improvement over GLiNER-x-base and matches the performance of much larger models such as Qwen3-32B, while offering significantly higher inference efficiency.

model designmultilingual NERnamed entity recognition

FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition

Dec 15, 2025
JG
Jonas Golde
🏛️ Humboldt Universität zu Berlin

Current multilingual named entity recognition (NER) research is hindered by the lack of systematically constructed, reusable large-scale annotated datasets. To address this, we introduce the first scalable synthetic NER dataset covering 91 languages and 25 scripts. Our method leverages FineWeb-Edu to filter NER-relevant texts, employs multilingual large language models (LLMs) for automated annotation, and integrates LLM-as-a-judge for quality assessment. We further propose a novel regression-model pre-screening mechanism to enhance annotation fidelity (3.99/5) and completeness (4.05/5). The dataset comprises 225K text segments and 235K entities, featuring bilingual (source + English) standardized labels—a first in multilingual NER. Using a teacher-student paradigm with cross-lingual label alignment, our approach achieves zero-shot transfer performance on English, Thai, and Swahili that matches or surpasses baselines using only 1/19 of their training data; the regression model attains an F1 score of 84.1%. The full dataset and toolchain are open-sourced.

Creating scalable multilingual NER datasets via teacher-student paradigmImproving zero-shot transfer performance with less training dataProviding high-quality annotations with translated labels for diverse languages

Hot Scholars

DR

Davide Rigoni

University of Padova
Artificial IntelligenceMachine LearningDeep LearningGenerative Models
JL

Jiaying Lu

Research Assistant Professor of School of Nursing's Center for Data Science, at Emory University
AI for HealthcareKnowledge GraphMultimodal LearningLarge Language Model
AJ

Alípio Jorge

University of Porto, FCUP, DCC, INESC TEC, LIAAD
Machine LearningNLPNarrative ExtractionRecommender Systems
CC

Cornelia Caragea

University of Illinois at Chicago
Natural Language ProcessingDeep LearningInformation RetrievalArtificial Intelligence