Score
Designs and implements systems and pipelines that detect and extract entity mentions, classify and normalize their surface forms, and resolve duplicates or coreference to produce consolidated entity records. Builds and analyzes algorithms and workflows for disambiguation and linking — including batched, two-phase/two-stage, oracle-assisted, cross-database and cross-lingual strategies — that map mentions to canonical identifiers (e.g., language-independent knowledge-base IDs) and merge clusterings into a unified linked-entity inventory.
This paper addresses heterogeneous entity matching (HEM), a core challenge in data integration, arising from structural, syntactic, schema, and semantic heterogeneity. To tackle the resulting modeling difficulties, we propose the first unified classification framework that jointly accounts for representational and semantic heterogeneity, grounded in the FAIR principles to expose fundamental limitations of existing methods under semantic inconsistency. Through a systematic literature review, taxonomy-driven modeling, and cross-model experimental evaluation, we empirically demonstrate the robustness deficiencies of mainstream entity matching models in semantically heterogeneous settings. Our analysis identifies key research directions—including multimodal fusion, human-in-the-loop approaches, joint modeling with large language models and knowledge graphs—as critical for advancing HEM. The work establishes a theoretical foundation, provides a standardized evaluation benchmark, and outlines a principled technical roadmap for future HEM research.
Traditional entity linking adopts a two-stage paradigm (mention detection followed by disambiguation), suffering from error propagation, high computational overhead, and poor cross-domain generalization. This paper proposes an end-to-end joint modeling framework that unifies mention detection and entity disambiguation. It leverages fine-tuned large language models (LLMs) to construct context-aware mention representations, effectively mitigating domain shift. Key contributions include: (i) a lightweight adapter mechanism that fuses deep semantic features from LLMs with local structural cues—avoiding full-parameter fine-tuning; and (ii) a cross-domain robust joint optimization objective. The method achieves state-of-the-art performance on multiple standard benchmarks (AIDA, MSNBC, ACE2005), notably improving F1 scores by an average of +4.2% in zero-shot cross-domain settings, thereby demonstrating superior effectiveness and generalization capability.
This work addresses data contamination caused by irreversible entity merging and ontology misclassification based on name fragments in knowledge graph construction. The authors propose a “review-before-linking” mechanism featuring an identity-ladder strategy—leveraging identifiers, names, and type scopes—to enable controlled deduplication, alongside anchor-evidence constraints that govern multi-class ontology label assignment. This approach corrects the evidential asymmetry arising when names are treated as instance labels rather than type assertions. Integrated into a system combining automated merging, evidence validation, and a human review queue, the method was evaluated on a knowledge graph comprising 537,157 entities and 2,198,567 relations. It reduced role assignment errors from 36 to zero, requiring only 775 manual decisions to resolve 48,403 merge proposals, thereby significantly mitigating risks of over-merging and misclassification.
This work addresses the poor generalizability of existing self-service entity resolution methods on unseen datasets and their difficulty in balancing precision and recall, which often leads to error propagation and cascading erroneous merges. To overcome these limitations, the authors propose a robust self-service entity resolution system that integrates three key innovations: automated selection among multiple algorithms, decoupled optimization of precision and recall—enforcing precision through rule-based veto mechanisms while enhancing recall via diverse candidate generation—and non-transitive cross-cluster merge validation to prevent error propagation. Extensive experiments on six real-world benchmark datasets, ranging from 864 to 5 million records, demonstrate that the proposed approach significantly reduces incorrect merges, enhances practical utility, and substantially lowers the trial-and-error burden for practitioners.
Traditional entity resolution (ER) relies on costly pairwise comparisons, limiting scalability; existing large language model (LLM)-based approaches remain confined to pairwise matching and fail to harness LLMs’ end-to-end clustering capability. This paper proposes LLM-CER, the first LLM-based ER framework that enables record-level direct clustering via in-context learning (ICL), eliminating pairwise comparisons entirely. By systematically exploring the clustering design space—including cluster size, diversity, and ordering—and integrating diversity-aware sampling, dynamic cluster merging, and hallucination-mitigating prompt engineering, LLM-CER jointly optimizes accuracy, efficiency, and robustness. Evaluated on nine real-world datasets, it achieves a 10% improvement in F1-score over the strongest baseline, up to a 150% gain in clustering accuracy, a fivefold reduction in API calls, and comparable monetary cost.
To address progressive entity resolution under time-sensitive and resource-constrained settings, this paper proposes the first systematic progressive entity matching framework, decomposing the process into four stages: filtering, weighting, scheduling, and matching. It introduces a unified design space encompassing mainstream approaches and a novel confidence-prioritized dynamic scheduling mechanism, significantly improving early recall and response efficiency. The framework integrates hybrid (rule- and learning-based) filtering, multi-feature weighting, and configurable matching functions, enabling end-to-end, on-demand retrieval of high-confidence results. Extensive experiments across 10 linkage and 8 deduplication datasets demonstrate that the optimal configuration achieves 90% recall of true duplicate pairs 42% earlier on average, while simultaneously improving both F1 score and time efficiency.
Traditional knowledge bases derived from large language models often suffer from ambiguity and unreliability due to their reliance on surface-level string matching, which fails to disambiguate homonyms or consolidate synonymous expressions. This work proposes a recursive knowledge extraction framework that incorporates a context-guided entity disambiguation mechanism during construction, enabling, for the first time in LLM-derived knowledge bases, effective synonym consolidation and homonym separation. The resulting knowledge base comprises 38.4 million triples, 1.6 million canonicalized entities, 207,600 integrated relations, and 66,000 unified categories. It further integrates entity linking, relation and category clustering, a SPARQL query engine, and a natural language-to-SPARQL translation module, offering an auditable, browsable, and queryable interactive web platform alongside full public data release.
Information extraction (IE) outputs often mismatch downstream database schemas, hindering direct integration. Method: This paper introduces TEXT2DB—a novel task requiring models to dynamically perform data completion, row insertion, and column expansion based on user instructions, document collections, and target database schemas. To address it, we propose OPAL, an agent framework operating via an Observe-Plan-Analyze closed-loop that orchestrates database interaction, code generation, IE model invocation, and pre-execution feedback analysis for end-to-end instruction understanding, schema alignment, and structured data population. Contribution/Results: Experiments demonstrate that OPAL accurately executes complex IE-database joint tasks across diverse database systems, significantly improving information-to-database deployment efficiency. The study further identifies critical challenges—including large-scale schema adaptation and model hallucination—highlighting open research directions for robust database-grounded IE.
This work addresses the challenge of entity resolution under strict batch query constraints, where dataset sizes far exceed per-query batch limits and individual batches cannot guarantee inclusion of all records pertaining to the same entity. The paper formally defines this batch-constrained entity resolution problem for the first time, proves that optimal batch selection is NP-hard, and proposes an efficient algorithm under natural assumptions on entity size distributions. By integrating combinatorial optimization with clustering theory, the method devises an adaptive querying strategy that dynamically constructs batches using prior knowledge of entity sizes, thereby controlling cost on a pay-as-you-go basis while maximizing recall at each step. Experiments on six real-world datasets demonstrate that the approach significantly outperforms state-of-the-art baselines, achieving higher recall under stringent query budget constraints.
This work addresses the inconsistency of cross-document software mentions in scientific literature by proposing a normalization framework that integrates semantic embeddings, knowledge base retrieval, and density-based clustering. The approach combines Sentence-BERT for semantic representation with FAISS for efficient retrieval, and incorporates surface form normalization and abbreviation resolution to handle out-of-vocabulary and ambiguous mentions. To enhance scalability in large-scale settings, an entity-type-aware blocking strategy is introduced. Guided by semantic centroids, HDBSCAN clustering enables the system to achieve CoNLL F1 scores of 0.98, 0.98, and 0.96 on the three subtasks of the SOMD 2026 shared task, substantially outperforming baseline methods.
This paper addresses key challenges in cross-source entity relationship modeling—namely, the difficulty of automatically identifying explicit/implicit dependencies, excessive manual intervention, and weak support for multilingual and heterogeneous environments. To this end, we propose the first end-to-end, data-stream-oriented entity relationship modeling framework. Our method integrates dynamic sampling, robust data analysis, SQL parsing, and natural language interfaces to automatically discover database constraints and construct both explicit and implicit dependency models, supporting real-time processing and visualization across major SQL dialects. Key contributions include: (i) the first data-stream-driven relational modeling approach that jointly optimizes semantic understanding and system efficiency; and (ii) multi-backend (CPU/GPU) and multilingual deployment capability. Evaluated on the STATS benchmark, our framework achieves 2.4× higher distributional representation efficiency, 2.6× faster constraint learning speed, 2.15× greater inference throughput, 1.19× improved data narrative accuracy, and 1.86× reduced contextual resource consumption.