In-context Clustering-based Entity Resolution with Large Language Models: A Design Space Exploration

📅 2025-06-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional entity resolution (ER) relies on costly pairwise comparisons, limiting scalability; existing large language model (LLM)-based approaches remain confined to pairwise matching and fail to harness LLMs’ end-to-end clustering capability. This paper proposes LLM-CER, the first LLM-based ER framework that enables record-level direct clustering via in-context learning (ICL), eliminating pairwise comparisons entirely. By systematically exploring the clustering design space—including cluster size, diversity, and ordering—and integrating diversity-aware sampling, dynamic cluster merging, and hallucination-mitigating prompt engineering, LLM-CER jointly optimizes accuracy, efficiency, and robustness. Evaluated on nine real-world datasets, it achieves a 10% improvement in F1-score over the strongest baseline, up to a 150% gain in clustering accuracy, a fivefold reduction in API calls, and comparable monetary cost.

Technology Category

Machine Learning: ClusteringNatural Language Processing: (Large) Language ModelsData Mining & Knowledge Management: Linked Open Data, Knowledge Graphs & KB Completion

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Large language models for searchUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Entity Resolution (ER) is a fundamental data quality improvement task that identifies and links records referring to the same real-world entity. Traditional ER approaches often rely on pairwise comparisons, which can be costly in terms of time and monetary resources, especially with large datasets. Recently, Large Language Models (LLMs) have shown promising results in ER tasks. However, existing methods typically focus on pairwise matching, missing the potential of LLMs to perform clustering directly in a more cost-effective and scalable manner. In this paper, we propose a novel in-context clustering approach for ER, where LLMs are used to cluster records directly, reducing both time complexity and monetary costs. We systematically investigate the design space for in-context clustering, analyzing the impact of factors such as set size, diversity, variation, and ordering of records on clustering performance. Based on these insights, we develop LLM-CER (LLM-powered Clustering-based ER), which achieves high-quality ER results while minimizing LLM API calls. Our approach addresses key challenges, including efficient cluster merging and LLM hallucination, providing a scalable and effective solution for ER. Extensive experiments on nine real-world datasets demonstrate that our method significantly improves result quality, achieving up to 150% higher accuracy, 10% increase in the F-measure, and reducing API calls by up to 5 times, while maintaining comparable monetary cost to the most cost-effective baseline.
Problem

Research questions and friction points this paper is trying to address.

Reducing time and cost in Entity Resolution using LLMs
Exploring LLM-based clustering for scalable Entity Resolution
Improving accuracy and efficiency in record clustering with LLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-context clustering with LLMs for ER
Reduces time complexity and API calls
Addresses cluster merging and LLM hallucination
J
Jiajie Fu
Zhejiang University
H
Haitong Tang
Zhejiang University
A
Arijit Khan
Aalborg University
S
S. Mehrotra
University of California, Irvine
X
Xiangyu Ke
Zhejiang University
Yunjun Gao
Yunjun Gao
Professor of Computer Science, Zhejiang University
DatabaseBig Data Management and Analyticsand AI Interaction with DB Technology