SCoRE: Streamlined Corpus-based Relation Extraction using Multi-Label Contrastive Learning and Bayesian kNN

๐Ÿ“… 2025-07-09
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
To address the challenges of poor robustness, high energy consumption, and difficulty in adapting large language models (LLMs) in low-supervision relation extraction (RE), this paper proposes a fine-tuning-free, modular, and energy-efficient RE framework. Methodologically, it innovatively integrates supervised multi-label contrastive learning with a Bayesian k-nearest neighbors (kNN) classifier to effectively mitigate noise inherent in distant supervision. It further introduces two fine-grained evaluation metricsโ€”Class-wise Semantic Discriminability (CSD) and Precision-at-R (P@R)โ€”and releases Wiki20d, a benchmark dataset designed to reflect realistic deployment scenarios. Experiments demonstrate that our approach achieves or surpasses state-of-the-art performance across five mainstream benchmarks while substantially reducing computational energy consumption. Ablation studies validate the advantages of its minimalist architecture in terms of robustness, zero-shot transferability, and seamless plug-and-play integration with LLMs, establishing an efficient and practical paradigm for automated knowledge graph expansion.

Technology Category

Data Mining & Knowledge Management: Linked Open Data, Knowledge Graphs & KB CompletionNatural Language Processing: Information ExtractionMachine Learning: Large Multimodal Models (LMMs)

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
๐Ÿ“ Abstract
The growing demand for efficient knowledge graph (KG) enrichment leveraging external corpora has intensified interest in relation extraction (RE), particularly under low-supervision settings. To address the need for adaptable and noise-resilient RE solutions that integrate seamlessly with pre-trained large language models (PLMs), we introduce SCoRE, a modular and cost-effective sentence-level RE system. SCoRE enables easy PLM switching, requires no finetuning, and adapts smoothly to diverse corpora and KGs. By combining supervised contrastive learning with a Bayesian k-Nearest Neighbors (kNN) classifier for multi-label classification, it delivers robust performance despite the noisy annotations of distantly supervised corpora. To improve RE evaluation, we propose two novel metrics: Correlation Structure Distance (CSD), measuring the alignment between learned relational patterns and KG structures, and Precision at R (P@R), assessing utility as a recommender system. We also release Wiki20d, a benchmark dataset replicating real-world RE conditions where only KG-derived annotations are available. Experiments on five benchmarks show that SCoRE matches or surpasses state-of-the-art methods while significantly reducing energy consumption. Further analyses reveal that increasing model complexity, as seen in prior work, degrades performance, highlighting the advantages of SCoRE's minimal design. Combining efficiency, modularity, and scalability, SCoRE stands as an optimal choice for real-world RE applications.
Problem

Research questions and friction points this paper is trying to address.

Efficient relation extraction for knowledge graph enrichment
Adaptable RE solutions with noise resilience
Improving evaluation metrics for relation extraction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-label contrastive learning for relation extraction
Bayesian kNN classifier for noisy annotations
Novel metrics CSD and P@R for evaluation
๐Ÿ’ผ Related Jobs
No related jobs found.
L
Luca Mariotti
Department of Physical, Computer and Mathematical Sciences - University of Modena and Reggio Emilia, via Giuseppe Campi, 213/a, Modena, 41125, Emilia Romagna, Italy
V
Veronica Guidetti
Department of Physical, Computer and Mathematical Sciences - University of Modena and Reggio Emilia, via Giuseppe Campi, 213/a, Modena, 41125, Emilia Romagna, Italy
Federica Mandreoli
Federica Mandreoli
FIM - University of Modena and Reggio Emilia - Italy
Database_systems