llm-driven ontology evolution

Designs and builds systems and pipelines that use large language models to automatically induce, extend, and refine ontological structures—class hierarchies, labels, and semantic relations—by extracting candidate concepts and hierarchical links from text or structured inputs. Implements automated validation, merging, and editable evolution workflows so ontologies can be iteratively updated, scaled across contexts, and maintained for downstream use.

llm-drivenontologyevolution

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.43
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Traditional ontology construction relies heavily on manual effort, suffering from poor scalability, low consistency, and limited adaptability. This work proposes a structured, iterative approach leveraging large language models (LLMs) to automate knowledge extraction and ontology component generation by integrating domain-specific context, while enabling continuous refinement. The method substantially accelerates the ontology development process, enhances semantic consistency, mitigates model bias, and improves transparency in the engineering workflow. Evaluation through a case study on constructing a user persona ontology in the automotive sales domain demonstrates that the proposed approach efficiently yields a highly consistent, scalable, and domain-specific knowledge base.

adaptabilityconsistencymanual development

Investigating Language Model Capabilities to Represent and Process Formal Knowledge: A Preliminary Study to Assist Ontology Engineering

Sep 12, 2025
HA
Hanna Abi Akl
🏛️ Université Côte d’Azur | Inria | CNRS | I3S | Data ScienceTech Institute (DSTI)

Large language models (LMs) exhibit weak reasoning capabilities and poor controllability in ontology engineering tasks. Method: This work investigates the feasibility of enhancing formal knowledge representation and reasoning performance in small language models (SLMs) by replacing natural language inputs with compact formal logic syntax—specifically, fragments of description logic—and conducts controlled experiments to systematically evaluate how syntactic formalization affects SLM performance on core ontology tasks, including consistency checking and classification. Contribution/Results: Formalized inputs significantly improve model interpretability and controllability while maintaining or even surpassing natural-language-based accuracy across multiple reasoning tasks. These findings establish a novel paradigm for trustworthy ontology construction that synergistically integrates symbolic logic with neural models, providing empirical validation for neuro-symbolic integration in knowledge engineering.

Assessing language models' formal knowledge representation capabilitiesExploring small language models for ontology engineering assistanceInvestigating compact logical languages for reasoning tasks

Assessing the Capability of Large Language Models for Domain-Specific Ontology Generation

Apr 24, 2025
AS
Anna Sofia Lippolis
🏛️ University of Bologna | ISTC-CNR | Linköping University

Traditional ontology engineering relies heavily on manual effort and suffers from poor generalizability across domains. Method: This study systematically evaluates the applicability and cross-domain generalization capability of large language models (LLMs) for automated ontology generation, proposing a capability-question (CQ)-driven prompt engineering framework. We empirically assess DeepSeek and o1-preview across six real-world ontology engineering projects using 95 curated CQs. Contribution/Results: We present the first empirical validation that structurally reasoning-capable LLMs can consistently generate high-quality ontologies across diverse domains—demonstrating domain-agnostic performance. Both models reliably perform ontology schema extraction, class/property identification, and logical axiom generation. This significantly enhances automation, scalability, and reproducibility in ontology construction. Our work establishes a novel paradigm for general-purpose ontology engineering, reducing reliance on domain-specific expertise and manual curation while maintaining formal expressivity and semantic fidelity.

Assess scalability of LLM-based ontology construction methodsEvaluate LLMs' ability to generate domain-specific ontologiesTest generalizability of DeepSeek and o1-preview models across domains

This study addresses the challenge of high costs associated with manual ontology construction in specialized domains, which often results in a lack of authoritative reference resources. It presents the first systematic exploration of leveraging large language models (LLMs) for automated domain ontology development. Focusing on Brazil’s “Blue Amazon” maritime territory as a case study, the authors employ prompt engineering to guide GPT-3.5 and GPT-4 in emulating domain experts, enabling the automatic generation of structured concept hierarchies from initial seed concepts. The experimental pipeline produced twenty ontologies, which expert evaluators deemed largely coherent and logically organized. Although minor human refinement remains necessary, the results strongly demonstrate the feasibility and practical potential of using LLMs as virtual experts to support ontology construction in knowledge-intensive domains.

Domain OntologyKnowledge RepresentationLarge Language Models

Large Language Models for Scholarly Ontology Generation: An Extensive Analysis in the Engineering Field

Dec 11, 2024
TA
Tanay Aggarwal
🏛️ The Open University | University of Milano Bicocca

To address high annotation costs, delayed updates, and poor generalizability in academic ontology construction, this paper systematically investigates the capability of large language models (LLMs) to automatically extract semantic relations—such as hypernymy/hyponymy and equivalence—among engineering research topics. Leveraging the IEEE Thesaurus, we construct a high-quality gold-standard dataset and conduct the first zero-shot prompting-based comparative evaluation across 17 LLMs, systematically varying model scale, openness (open vs. closed), and quantization level. Results show that lightweight quantized models—e.g., Dolphin-Mistral-7B—achieve an F1 score of 0.920 after prompt optimization, approaching the performance of the state-of-the-art closed-source model Claude 3 Sonnet (0.967). This demonstrates a significant breakthrough in balancing computational efficiency and accuracy, enabling resource-efficient, dynamic, and scalable academic knowledge structuring.

Automating scholarly ontology generation to replace manual creationEvaluating LLMs in identifying semantic relationships between research topicsOptimizing smaller models to match large proprietary models' performance

Latest Papers

What's happening recently
View more

This work addresses the limitations of large language models (LLMs) in long-term memory retention, structured understanding, and multi-step reasoning by proposing a hybrid intelligent system architecture. The approach employs an automated pipeline to construct RDF/OWL ontologies from heterogeneous data as an external memory layer, integrating vector-based retrieval with graph-based reasoning to establish an LLM-driven generate–verify–refine loop. By combining named entity recognition, relation extraction, triple generation, and SHACL/OWL constraint validation, the system substantially enhances the explainability and reliability of its inferences. Evaluated on multi-step planning tasks such as the Tower of Hanoi, the proposed framework outperforms baseline LLMs and enables formal verification and systematic error correction of its outputs.

explainabilityknowledge persistencelong-term memory

Traditional ontology extension methods are resource-intensive and error-prone, while existing large language model–based approaches lack explicit alignment with user requirements and reusable core ontologies, and suffer from insufficient systematic evaluation. This work proposes the first ontology extension framework that integrates competency question–driven design with retrieval-augmented generation (RAG), enabling context-aware, requirement-guided generation of ontology fragments by explicitly linking user needs—formulated as competency questions—with existing ontological knowledge. Evaluated on two real-world use cases, the generated fragments exhibit sound structural integrity and pass all functional tests; expert engineers assessed them as requiring only minor to moderate revisions for integration. These results demonstrate the feasibility, scalability, and evalability of the proposed approach.

competency questionsLarge Language Modelsontology extension

This study addresses the fragmentation in ontology learning caused by the absence of a unified infrastructure, which has led to disjointed methodologies, domain-specific approaches, and inconsistent evaluation practices. To overcome these challenges, this work proposes a modular, cross-domain ontology learning framework that integrates ontology access, a large language model (LLM)-driven learning pipeline, and standardized benchmarking. The project establishes the first comprehensive resource comprising 180 machine-readable ontologies spanning 22 domains and introduces standardized datasets for three core ontology learning tasks: term typing, taxonomy discovery, and non-taxonomic relation extraction. A collaborative approach combining LLMs and retrieval models is employed to support these tasks. Large-scale experiments reveal that performance bottlenecks stem primarily from mismatches between knowledge encoding and ontology structure rather than model capacity, while also demonstrating the efficacy of cross-domain, multi-task evaluation as a foundation for systematic ontology learning research.

BenchmarkingKnowledge RepresentationLarge Language Models

This work addresses the challenge that large language models (LLMs) struggle to adhere to formal semantic constraints in real time when generating structured knowledge, often relying on inefficient and error-prone post-hoc validation. To overcome this limitation, the authors propose an ontology-to-tool compilation mechanism that automatically translates domain ontology specifications into executable tool interfaces. By compelling LLM agents to interact with knowledge graphs exclusively through these generated tools, the approach proactively enforces semantic consistency during knowledge generation. Built upon The World Avatar framework, the method integrates the Model Context Protocol, ontology-driven tool synthesis, and agent workflows, substantially reducing the need for manual prompt engineering. Evaluated on the task of processing scientific literature on metal–organic polyhedra synthesis, the system successfully guides LLMs to extract, validate, and repair structured knowledge, demonstrating the feasibility and advantages of this paradigm for scientific text understanding.

executable semanticsknowledge graphlarge language models

Automatically generating high-quality formal ontologies from unstructured text remains challenging, as existing large language model (LLM)-based approaches are often hindered by ambiguous design, structural redundancy, and ineffective repair mechanisms. This work proposes a planning-first, artifact-driven multi-agent paradigm for ontology generation, decomposing the task into a collaborative workflow among four specialized roles: domain expert, manager, coder, and quality assurer. The framework integrates heterogeneous LLM-based review, SPARQL competency assessment, and retrieval-augmented generation to iteratively refine ontological artifacts. Compared to single-agent baselines, the proposed approach substantially improves the structural quality and auditability of generated ontologies while moderately enhancing their query usability.

knowledge engineeringlarge language modelsontology design patterns

Hot Scholars

AG

Adrian Groza

Technical University of Cluj-Napoca, European University of Technology (EUt+)
Artificial IntelligenceAgentic AIKnowledge representationExplainable AI
AL

Alexandru Lecu

Technical University of Cluj-Napoca, Digital Science, European University of Technology (EUt+)
Natural Language ProcessingArtificial IntelligenceKnowledge RepresentationLLMs
ND

Neil De La Fuente

Student Researcher, Technical University of Munich
Deep LearningSynthetic DataComputer VisionSelf Supervised Learning
AJ

Alekh Jindal

CEO and Co-founder, Tursio Inc.
Database SystemsInformation SystemsCloud Computing
QZ

Qiang Zhang

Zhejiang University
Machine learningAI for Science