domain knowledge graph construction

Designs and builds structured, queryable knowledge graphs and ontologies that formalize domain-specific entities, hierarchical concepts, and relationships by extracting and modeling information from documents, specifications, and other sources. Implements end-to-end extraction and integration pipelines — parsers and document-level extractors, schema and hierarchy modeling, alignment and grounding to canonical vocabularies, mapping and merging of taxonomies, and graph-database ingestion — and analyzes the resulting graphs for consistency, completeness, and queryability.

domainknowledgegraphconstruction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
2.66
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$187K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of traditional knowledge graph construction approaches, wherein structural decisions are hard-coded into rigid pipelines, resulting in tight coupling between schema and construction process and hindering support for ontology-level tasks. To overcome this, the authors propose an ontology-oriented construction framework featuring a novel intrinsic-relational routing mechanism. This mechanism dynamically assigns attributes to corresponding schema modules through iterative attribute classification, enabling a declarative and reusable decoupled design. The pipeline integrates rule-based cleaning, tool-augmented large language model–assisted annotation, and human review. Evaluated on Wikidata (January 2026), the resulting graph comprises 34 million nodes and 61.2 million edges, achieving 93.3% schema coverage and 98.0% module assignment accuracy, effectively supporting five ontology-level applications.

knowledge graphontologyproperty graph

Graph Queries from Natural Language using Constrained Language Models and Visual Editing

Nov 30, 2025
BK
Benedikt Kantz
🏛️ Graz University of Technology

To address the challenge non-expert users face in directly querying knowledge graphs, this paper proposes a natural language-driven interactive query construction method. The approach employs a two-stage constrained language model that integrates ontology-based semantic constraints to generate syntactically and semantically valid query prototypes—thereby avoiding invalid classes, relations, and grammatical errors. A visual editor enables users to iteratively refine queries via natural language descriptions and graphical adjustments. Finally, an interpretable SPARQL translation pipeline converts the refined prototype into standard SPARQL. Evaluated across multiple ontologies and language models, the system consistently produces correct SPARQL queries without manual intervention, outperforming existing baselines in both retrieval accuracy and efficiency. Validation through synthetic data experiments and an initial user study confirms the method’s effectiveness, usability, and practical applicability.

Consistently generates valid SPARQL queries without syntax or ontology correctionsConverts natural language to prototype graphs using constrained language model generationEnables non-experts to query knowledge graphs via natural language and visual editing

Domain specific ontologies from Linked Open Data (LOD)

Jan 08, 2022
RA
Rosario A. Uceda-Sosa
🏛️ IBM

Entity disambiguation and linking in IT domains suffer from poor domain adaptability and difficulty in incorporating domain-specific knowledge when relying solely on general-purpose knowledge graphs (e.g., Wikidata, DBpedia). Method: This paper proposes a lightweight, extensible ontology construction method tailored for the IT domain. Starting from general Linked Open Data (LOD) resources, it employs a domain-agnostic pipeline and—novelly—integrates an IT-specific terminology lexicon to drive ontology schema expansion. The approach synergistically combines SPARQL querying, RDF reasoning, and ontology alignment. Contribution/Results: The resulting paradigm balances generality and domain specificity, significantly improving accuracy in IT entity disambiguation and linking. It establishes a low-barrier, reusable ontology engineering framework that supports continuous injection of proprietary domain knowledge, thereby enabling sustainable, scalable domain ontology development.

Bootstrapping IT ontology using domain-agnostic and domain-specific methodsEnhancing entity disambiguation with domain-specific knowledge graphsImproving efficiency in consuming and extending proprietary content

Ontology-grounded Automatic Knowledge Graph Construction by LLM under Wikidata schema

Dec 30, 2024
XF
Xiaohan Feng
🏛️ Chinese University of Hong Kong

To address low automation, poor interpretability, and semantic incompatibility with Wikidata in knowledge graph (KG) construction, this paper proposes an ontology-driven large language model (LLM) approach. First, domain scope and relations are lightweightly extracted from Competency Questions (CQs). Second, extracted relations are bidirectionally mapped to the Wikidata ontology to achieve semantic alignment. Third, LLMs are guided—under ontology constraints—to generate structured subject-predicate-object triples. The key contribution is the first CQ-driven ontology construction paradigm, uniquely balancing automation and interpretability. Experiments on standard benchmarks demonstrate state-of-the-art performance; the resulting KG exhibits high logical consistency and cross-system interoperability, significantly reducing reliance on manual annotation. The method enables a scalable, reusable KG construction pipeline.

Knowledge Graph ConstructionLarge-scale Language ModelsWikidata Compatibility

Semantic Web: Past, Present, and Future

Dec 22, 2024
AS
A. Scherp
🏛️ Ulm University | Carl Zeiss SMT GmbH | Charles University | TU Wien | Leibniz University of Hannover | TIB-Leibniz Information Centre for Science and Technology

This paper addresses two critical gaps in the evolution of the Semantic Web: (1) the theoretical lag behind practical applications, and (2) insufficient integration of trustworthiness mechanisms with AI. To bridge these gaps, we propose a unified analytical framework that synergistically integrates classical semantic technologies with modern AI. Methodologically, we extend the canonical “layered cake” model into a novel three-dimensional paradigm encompassing trustworthy computing, industrial validation, and LLM–KG co-adaptation—systematically unifying RDF/OWL representation, rule-based reasoning, distributed SPARQL query processing, knowledge graph embedding, graph neural networks, and LLM–KG alignment techniques. Our contributions include: (1) a comprehensive technology landscape charting 50 years of Semantic Web development; (2) a clarified integration roadmap for knowledge graphs and AI—particularly large language models; and (3) theoretical foundations and practical guidelines for building next-generation semantic infrastructure that is trustworthy, interpretable, and adaptive.

Enhance traditional Semantic Web concepts with recent developments like provenanceExplore machine learning methods on knowledge graphs and language model relationsRecap classical foundations and modern applications of Semantic Web technologies

Latest Papers

What's happening recently
View more

This work addresses data contamination caused by irreversible entity merging and ontology misclassification based on name fragments in knowledge graph construction. The authors propose a “review-before-linking” mechanism featuring an identity-ladder strategy—leveraging identifiers, names, and type scopes—to enable controlled deduplication, alongside anchor-evidence constraints that govern multi-class ontology label assignment. This approach corrects the evidential asymmetry arising when names are treated as instance labels rather than type assertions. Integrated into a system combining automated merging, evidence validation, and a human review queue, the method was evaluated on a knowledge graph comprising 537,157 entities and 2,198,567 relations. It reduced role assignment errors from 36 to zero, requiring only 775 manual decisions to resolve 48,403 merge proposals, thereby significantly mitigating risks of over-merging and misclassification.

data curationentity identityentity resolution

This work addresses the need for unified and efficient knowledge provisioning in large language models by proposing a novel architecture that integrates relational and property graph data models. The approach leverages record addresses from log files as immutable reference values in place of traditional foreign keys, enabling efficient graph-style link traversal instead of costly join queries while natively supporting triple-based knowledge representation. The resulting unified knowledge service framework combines the structural rigor of relational models with the flexible associative capabilities of graph models, significantly enhancing knowledge retrieval efficiency and effectively supporting knowledge integration and invocation in generative AI systems.

data integrationknowledge representationlarge language models

This work addresses the challenge of efficiently and accurately extracting complex nested structures from unstructured text and enabling semantic-level automated evaluation. The authors propose a schema-guided, end-to-end framework that integrates domain-specific knowledge schemas with generative AI models—such as Claude Opus 3—to perform zero-shot, one-pass extraction of hierarchical attributes with variable cardinality. Automated evaluation is achieved through path alignment and fine-grained semantic matching algorithms. The framework demonstrates strong transferability across models, institutions, and languages, successfully extracting 12 out of 14 attributes in NICE documents with F1 scores exceeding 90%. It achieves a 30-fold speedup over manual annotation while significantly improving extraction efficiency, consistency, and generalizability.

generative AIhierarchical informationinformation extraction

This work addresses the limitations of existing ontology documentation tools in supporting modular modeling and human readability, particularly in handling cross-module entities and annotations. To overcome these challenges, the authors refactor and extend the LODE framework by introducing a modular Reader-Model-Viewer architecture that decouples parsing, modeling, and rendering components. Implemented as a web service, the new framework provides enhanced capabilities for generating OWL ontology documentation, featuring dedicated entity pages, RDF provenance tracking, and Markdown-based rendering. These improvements significantly increase the intelligibility and reusability of modular scientific knowledge graph ontologies. The framework has been successfully applied to the documentation of the SKG-O ontology, demonstrating its practical utility and effectiveness.

human-readable documentationmodular ontologiesontology documentation

This work addresses the limitations of large language models (LLMs) in long-term memory retention, structured understanding, and multi-step reasoning by proposing a hybrid intelligent system architecture. The approach employs an automated pipeline to construct RDF/OWL ontologies from heterogeneous data as an external memory layer, integrating vector-based retrieval with graph-based reasoning to establish an LLM-driven generate–verify–refine loop. By combining named entity recognition, relation extraction, triple generation, and SHACL/OWL constraint validation, the system substantially enhances the explainability and reliability of its inferences. Evaluated on multi-step planning tasks such as the Tower of Hanoi, the proposed framework outperforms baseline LLMs and enables formal verification and systematic error correction of its outputs.

explainabilityknowledge persistencelong-term memory

Hot Scholars

SA

Sören Auer

Leibniz University of Hannover, Leibniz TIB, L3S Research Center
Neurosymbolic AIKnowledge GraphsWeb ScienceDigital Libraries
JC

Jiaoyan Chen

Department of Computer Science, University of Manchester
Knowledge GraphOntologyMachine LearningLarge Language Model
EK

Evgeny Kharlamov

Bosch Center for Artificial Intelligence and Univeristy of Oslo
Nuero-Symbolic AIKnowledge GraphsAgentsRAG
AB

Andrea Bartolini

Associate Professor, University of Bologna
Energy managementThermal managementNear-Threshold ComputingHigh Performance Computing
YI

Yusuke Iwasawa

The University of Tokyo
deep learningtransfer learningfoundation modelmeta learning