Score
Designs, implements, and evaluates algorithms and systems that identify, parse, disambiguate, normalize, link, and merge mentions and records of real-world entities across texts, databases, and knowledge bases, producing reconciled canonical entities and cross-entity correlations. This work also covers entity-relationship modeling, knowledge extraction and modeling, knowledge-base integration and fusion, and the application of inference/reasoning to enrich, validate, and maintain integrated entity representations.
This paper addresses heterogeneous entity matching (HEM), a core challenge in data integration, arising from structural, syntactic, schema, and semantic heterogeneity. To tackle the resulting modeling difficulties, we propose the first unified classification framework that jointly accounts for representational and semantic heterogeneity, grounded in the FAIR principles to expose fundamental limitations of existing methods under semantic inconsistency. Through a systematic literature review, taxonomy-driven modeling, and cross-model experimental evaluation, we empirically demonstrate the robustness deficiencies of mainstream entity matching models in semantically heterogeneous settings. Our analysis identifies key research directions—including multimodal fusion, human-in-the-loop approaches, joint modeling with large language models and knowledge graphs—as critical for advancing HEM. The work establishes a theoretical foundation, provides a standardized evaluation benchmark, and outlines a principled technical roadmap for future HEM research.
Existing ASPEN systems support only global entity resolution, making them ill-suited for value-level heterogeneity—e.g., resolving “J. Lee” context-dependently to either “Joy Lee” or “Jake Lee”—and lack optimization criteria targeting parsing quality. This paper proposes ASPEN+, which addresses these limitations via three key innovations: (1) a local merging mechanism enabling context-sensitive clustering of identical entity names; (2) a novel optimization objective balancing conflict minimization and rule support maximization, formalized under multiple semantics of optimality; and (3) an extension of the original system using Answer Set Programming to enable fine-grained logical reasoning and efficient solving. Experiments on real-world datasets demonstrate that ASPEN+ significantly improves parsing accuracy while maintaining acceptable runtime efficiency.
This work addresses the need for unified and efficient knowledge provisioning in large language models by proposing a novel architecture that integrates relational and property graph data models. The approach leverages record addresses from log files as immutable reference values in place of traditional foreign keys, enabling efficient graph-style link traversal instead of costly join queries while natively supporting triple-based knowledge representation. The resulting unified knowledge service framework combines the structural rigor of relational models with the flexible associative capabilities of graph models, significantly enhancing knowledge retrieval efficiency and effectively supporting knowledge integration and invocation in generative AI systems.
This work addresses data contamination caused by irreversible entity merging and ontology misclassification based on name fragments in knowledge graph construction. The authors propose a “review-before-linking” mechanism featuring an identity-ladder strategy—leveraging identifiers, names, and type scopes—to enable controlled deduplication, alongside anchor-evidence constraints that govern multi-class ontology label assignment. This approach corrects the evidential asymmetry arising when names are treated as instance labels rather than type assertions. Integrated into a system combining automated merging, evidence validation, and a human review queue, the method was evaluated on a knowledge graph comprising 537,157 entities and 2,198,567 relations. It reduced role assignment errors from 36 to zero, requiring only 775 manual decisions to resolve 48,403 merge proposals, thereby significantly mitigating risks of over-merging and misclassification.
This work addresses the poor generalizability of existing self-service entity resolution methods on unseen datasets and their difficulty in balancing precision and recall, which often leads to error propagation and cascading erroneous merges. To overcome these limitations, the authors propose a robust self-service entity resolution system that integrates three key innovations: automated selection among multiple algorithms, decoupled optimization of precision and recall—enforcing precision through rule-based veto mechanisms while enhancing recall via diverse candidate generation—and non-transitive cross-cluster merge validation to prevent error propagation. Extensive experiments on six real-world benchmark datasets, ranging from 864 to 5 million records, demonstrate that the proposed approach significantly reduces incorrect merges, enhances practical utility, and substantially lowers the trial-and-error burden for practitioners.
This paper addresses two critical gaps in the evolution of the Semantic Web: (1) the theoretical lag behind practical applications, and (2) insufficient integration of trustworthiness mechanisms with AI. To bridge these gaps, we propose a unified analytical framework that synergistically integrates classical semantic technologies with modern AI. Methodologically, we extend the canonical “layered cake” model into a novel three-dimensional paradigm encompassing trustworthy computing, industrial validation, and LLM–KG co-adaptation—systematically unifying RDF/OWL representation, rule-based reasoning, distributed SPARQL query processing, knowledge graph embedding, graph neural networks, and LLM–KG alignment techniques. Our contributions include: (1) a comprehensive technology landscape charting 50 years of Semantic Web development; (2) a clarified integration roadmap for knowledge graphs and AI—particularly large language models; and (3) theoretical foundations and practical guidelines for building next-generation semantic infrastructure that is trustworthy, interpretable, and adaptive.
This study investigates the reliability and limitations of large language models (LLMs) in automatically generating entity-relationship (ER) diagrams from complex natural language requirements. Employing prompt strategies including zero-shot, chain-of-thought (CoT), and CoT augmented with a verifier, the authors systematically evaluate three leading LLMs on their ability to extract entities, relationships, and attributes from textual descriptions and produce conceptually consistent ER diagrams. The results indicate that while models perform adequately on low-complexity specifications, their outputs frequently suffer from logical inconsistencies, semantic ambiguities, and failures to correctly express constraints as requirement complexity increases. The findings highlight fundamental shortcomings of current LLMs in high-stakes database modeling tasks and provide empirical evidence for the role of prompt engineering in structured conceptual modeling.
This work addresses the limitations of traditional entity alignment methods under noisy or weakly supervised settings and the poor interpretability and high inference costs of existing large language model (LLM)-based approaches. It introduces, for the first time, a structured multi-step reasoning agent into the entity alignment task. The proposed framework employs a plan-and-execute mechanism to yield interpretable alignment decisions and incorporates attribute- and relation-based triple selectors to pre-filter redundant information, thereby enhancing computational efficiency. By integrating LLMs, multi-step reasoning, triple selection, and knowledge graph representation learning, the method achieves state-of-the-art performance across three benchmark datasets while simultaneously ensuring interpretability and computational efficiency, thus overcoming the black-box limitations commonly associated with LLM applications.
Traditional knowledge bases derived from large language models often suffer from ambiguity and unreliability due to their reliance on surface-level string matching, which fails to disambiguate homonyms or consolidate synonymous expressions. This work proposes a recursive knowledge extraction framework that incorporates a context-guided entity disambiguation mechanism during construction, enabling, for the first time in LLM-derived knowledge bases, effective synonym consolidation and homonym separation. The resulting knowledge base comprises 38.4 million triples, 1.6 million canonicalized entities, 207,600 integrated relations, and 66,000 unified categories. It further integrates entity linking, relation and category clustering, a SPARQL query engine, and a natural language-to-SPARQL translation module, offering an auditable, browsable, and queryable interactive web platform alongside full public data release.
Traditional entity resolution approaches struggle with uncertainty and context dependence in dynamic, heterogeneous data and lack active reasoning and interactive capabilities. This work proposes a novel paradigm—active entity resolution—reformulating the resolution process as a sequential decision-making problem for an autonomous agent. It integrates active planning, cross-source evidence reasoning, and human-in-the-loop collaboration. We formalize this problem, present the first reference architecture supporting external querying and proactive decision-making, and define core challenges alongside new evaluation metrics. By bridging data management and autonomous agent technologies, our approach enables more accurate, efficient, and cost-aware entity resolution.