Score
Designs, implements, and evaluates systems that detect and classify contiguous spans of natural language text corresponding to real-world entities (for example, persons, organizations, locations), producing labeled entity spans. This work covers tag/schema design, annotation and data pipelines, token- or span-level model architectures, and evaluation metrics for entity identification and classification.
This study addresses a critical gap in the literature by systematically evaluating the capacity of large language models (LLMs) to identify subjective text segments—a capability not previously assessed in a comprehensive manner. The work presents the first thorough examination of LLM performance across three representative tasks: sentiment analysis, offensive language detection, and claim verification. Leveraging a combination of instruction tuning, in-context learning, and chain-of-thought reasoning, the experiments demonstrate that LLMs effectively model intra-textual semantic relationships, substantially improving accuracy in segment-level subjectivity recognition. The findings underscore the pivotal role of leveraging textual structural information for fine-grained subjective understanding and establish both a methodological framework and empirical foundation for future research in this domain.
This work addresses key challenges in named entity recognition (NER) for 18th-century French historical encyclopedias—including nonstandard orthography, nested/overlapping entities, and severe annotation scarcity. To tackle these, we propose a novel dual-granularity classification framework that jointly models token-level and span-level predictions, reformulating NER as a unified token-plus-span co-identification task—the first such formulation. Methodologically, we integrate conditional random fields (CRF), spaCy, CamemBERT, Flair, and few-shot prompting of generative large language models, establishing a hybrid paradigm that synergizes symbolic rules with neural modeling. On the GeoEDdA benchmark, our Transformer-based model achieves an F1 score of 89.2% on nested entities; remarkably, the generative variant attains 62.4% F1 using only five annotated examples, demonstrating strong efficacy and generalization capacity in low-resource historical text processing.
This work addresses the high computational overhead and latency of existing span-based named entity recognition methods, which enumerate numerous candidate spans and rely on token-level augmentation, rendering them inefficient for industrial applications demanding real-time performance. To overcome these limitations, the authors propose SpanDec, a framework that concentrates span representation interaction exclusively in the final Transformer layer and introduces a lightweight decoder coupled with a dynamic candidate filtering mechanism. This design enables early pruning of low-quality spans, thereby eliminating redundant computations in earlier layers. SpanDec achieves accuracy comparable to state-of-the-art span-based models while significantly improving throughput and reducing computational cost, making it well-suited for high-concurrency services and edge deployment.
Large language models lack explicit mechanisms to refer to specific spans in the input text, leading to inconsistent performance with existing span annotation prompting strategies. This work systematically examines three categories of approaches: input tagging, numerical indexing, and content matching, and proposes LogitMatch—a novel constrained decoding method that enforces alignment between model outputs and valid input spans in logit space to address the inconsistency inherent in content matching. Experiments across four diverse tasks demonstrate that LogitMatch significantly outperforms existing content matching methods and, in certain settings, surpasses other strategies, while also confirming that input tagging remains a robust baseline.
Existing evaluation benchmarks for vision-rich document (VrD) information extraction suffer from severe biases: coarse-grained annotations introduce spurious correlations between inputs and labels, inflating model performance estimates; moreover, holistic F1-based evaluation fails to assess robustness in realistic scenarios. Method: We propose an entity-centric evaluation paradigm and introduce EC-FUNSD—the first semantic-driven benchmark for entity recognition and linking in VrDs. Contribution/Results: EC-FUNSD innovatively decouples paragraph-block layout annotations from semantic entity definitions, incorporating cross-block entity linking, fine-grained relational understanding, and diverse document layouts. Experiments show that state-of-the-art vision-language pre-trained models exhibit substantial performance degradation on EC-FUNSD, validating the overfitting issue inherent in prior benchmarks. EC-FUNSD thus establishes a more rigorous, semantically grounded, and robust evaluation standard for VrD understanding.
Existing natural language processing resources often lack task-specific information for niche or emerging entities, hindering accurate classification in domains such as business or healthcare provider categorization. To address this limitation, this work proposes a dynamic classification framework that requires no additional labeled text: given only entity names and their corresponding labels, the method retrieves web-based information and leverages large language models (LLMs) to generate task-relevant descriptions, which are then used to train a text classifier. This end-to-end approach achieves strong performance on low-resource entity classification, attaining macro-averaged F1 scores of 82.3% on Standard Industrial Classification (SIC) coding and 72.9% on healthcare provider categorization, thereby demonstrating its effectiveness and practical utility.
This work addresses data contamination caused by irreversible entity merging and ontology misclassification based on name fragments in knowledge graph construction. The authors propose a “review-before-linking” mechanism featuring an identity-ladder strategy—leveraging identifiers, names, and type scopes—to enable controlled deduplication, alongside anchor-evidence constraints that govern multi-class ontology label assignment. This approach corrects the evidential asymmetry arising when names are treated as instance labels rather than type assertions. Integrated into a system combining automated merging, evidence validation, and a human review queue, the method was evaluated on a knowledge graph comprising 537,157 entities and 2,198,567 relations. It reduced role assignment errors from 36 to zero, requiring only 775 manual decisions to resolve 48,403 merge proposals, thereby significantly mitigating risks of over-merging and misclassification.
This work addresses the challenges of multi-domain named entity recognition (NER) under conditions of scarce labeled data, where conventional approaches suffer from domain discrepancy, data sparsity, and overfitting. To overcome these limitations, the paper proposes a unified framework that integrates unsupervised pre-training, transfer learning, data augmentation, few-shot learning, and domain-adversarial training. This approach enables effective adaptation to target domains without requiring any annotated data therein, significantly enhancing model generalization and robustness in low-resource, multi-domain settings. Experimental results demonstrate that the proposed method substantially outperforms existing baselines, offering a novel and practical pathway toward efficient and transferable NER in resource-constrained environments.
This work addresses document-level conspiracy theory detection by proposing a joint framework that integrates multi-label span classification with sequence classification. For extracting conspiracy-related markers—such as roles and actions—the approach formulates the task as boundary-aware multi-label span classification, incorporating IoU-based positive labeling, hard negative sampling, and an inclusion-aware non-maximum suppression strategy, while distinguishing between entity-like and abstract roles. Document-level classification is performed using a RoBERTa model enhanced with label smoothing. Evaluated on SemEval-2026 Task 10, the method achieves 7th place in Subtask 1 (macro F1 = 0.2251) and 11th place in Subtask 2 (weighted F1 = 0.7694), demonstrating the effectiveness of the proposed techniques.