knowledge-base integration into models

Designs, builds, and evaluates systems that integrate structured domain knowledge—knowledge bases, ontologies, knowledge graphs, and rule sets—into machine learning and language models to enable knowledge-grounded generation, relation reasoning, and KG-augmented inference. This includes creating knowledge representations and semantic models, engineering rule-based and schema mappings between resources and model inputs, implementing knowledge-graph reasoning pipelines, and analyzing the correctness, fidelity, and interoperability of the integrated knowledge.

knowledge-baseintegrationintomodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.6
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$189K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey

May 06, 2025
DZ
Da Zheng
🏛️ Ant Group | The University of Hong Kong | Zhejiang University

This paper systematically investigates the capability boundaries and core challenges of large language models (LLMs) in solving complex, multi-step problems—specifically identifying three critical bottlenecks: weak multi-step reasoning, insufficient domain-knowledge integration, and unverifiable outputs. To address these, the authors introduce the first cross-domain evaluation framework spanning software engineering, mathematical proof, data analysis, and scientific discovery, and propose a “knowledge–reasoning–verification” co-evolution paradigm. Their method integrates chain-of-thought prompting, external knowledge retrieval, tool invocation, and multi-level verification to jointly inject structured and unstructured knowledge. Empirical analysis clarifies current LLM performance limits on complex tasks, yields a scalable knowledge-augmented architecture, and identifies six concrete research directions toward trustworthy AI—thereby providing both theoretical foundations and practical pathways for high-reliability AI decision-making.

Addressing challenges like multi-step reasoning and domain knowledge integrationExploring LLMs' capabilities and limitations in complex problem-solvingInvestigating verification techniques for LLM-based solutions across domains

Must-Read Papers

Most classic and influential ideas
View more

LLM-empowered knowledge graph construction: A survey

Oct 23, 2025
HB
Haonan Bian
🏛️ Xidian University

This paper addresses the limitations of traditional rule- and statistics-based approaches in knowledge graph (KG) construction—namely, ontology engineering, knowledge extraction, and knowledge fusion. It proposes a novel “language-driven generative KG construction” paradigm that unifies schema-based and schema-agnostic methods, establishing synergistic mechanisms between large language models (LLMs) and symbolic KGs for structured organization and open-ended semantic expression. The method integrates prompt engineering, knowledge representation learning, automated reasoning, and multimodal modeling to enable dynamic, interpretable knowledge acquisition and fusion. The study comprehensively surveys technical pathways and bottlenecks, identifying three key research directions: LLM reasoning enhancement via KGs, agent memory modeling with KGs, and multimodal KG construction. Ultimately, this work advances the development of adaptive, neuro-symbolic intelligent knowledge systems.

Analyzing LLM impact on ontology engineering and knowledge extractionBridging symbolic knowledge engineering with neural semantic understandingSurveying LLM-driven knowledge graph construction paradigms

This work proposes a novel knowledge graph representation paradigm that addresses the limitations of traditional approaches, which rely on predefined symbolic relations and consequently fail to capture the context-dependency, fine-grained semantics, and inherent uncertainty of real-world relationships—often leading to critical information loss. By reformulating relations as natural language descriptions rather than discrete symbols, the proposed framework leverages the generative and reasoning capabilities of large language models. It integrates prompt engineering with a minimal structural backbone, thereby harmonizing structured scaffolding with unstructured semantic expression. This hybrid design substantially enhances the semantic richness and context-awareness of relation modeling, offering a more adaptive and expressive pathway for knowledge graph construction in the era of large language models.

knowledge graphlarge language modelsnatural-language relations

Ontology-grounded Automatic Knowledge Graph Construction by LLM under Wikidata schema

Dec 30, 2024
XF
Xiaohan Feng
🏛️ Chinese University of Hong Kong

To address low automation, poor interpretability, and semantic incompatibility with Wikidata in knowledge graph (KG) construction, this paper proposes an ontology-driven large language model (LLM) approach. First, domain scope and relations are lightweightly extracted from Competency Questions (CQs). Second, extracted relations are bidirectionally mapped to the Wikidata ontology to achieve semantic alignment. Third, LLMs are guided—under ontology constraints—to generate structured subject-predicate-object triples. The key contribution is the first CQ-driven ontology construction paradigm, uniquely balancing automation and interpretability. Experiments on standard benchmarks demonstrate state-of-the-art performance; the resulting KG exhibits high logical consistency and cross-system interoperability, significantly reducing reliance on manual annotation. The method enables a scalable, reusable KG construction pipeline.

Knowledge Graph ConstructionLarge-scale Language ModelsWikidata Compatibility

On the Evolution of Knowledge Graphs: A Survey and Perspective

Oct 07, 2023
XJ
Xuhui Jiang
🏛️ International Digital Economy Academy | Chinese Academy of Sciences

Knowledge graphs (KGs) suffer from conceptual ambiguity across static, dynamic, temporal, and event-based variants, alongside limitations in KG construction and reasoning—particularly in handling multimodal data and temporal dynamics. Method: We propose a unified framework integrating symbolic rules with neural methods for knowledge extraction and reasoning, coupled with cross-modal alignment and temporal modeling. We establish the first lifecycle-spanning theoretical framework for KG evolution and introduce a novel KG–large language model (LLM) co-engineering paradigm. Contribution/Results: Our work rigorously delineates the conceptual boundaries and technical paradigms of four KG types; enables robust, scalable KG construction and temporal reasoning; and demonstrates empirical validity and extensibility through a financial risk identification application. The framework advances foundational KG theory and provides a systematic methodology for KG–LLM joint modeling, bridging symbolic and neural AI in knowledge-intensive domains.

Exploring knowledge extraction and reasoning techniquesProposing future directions for knowledge engineeringSurveying evolution of diverse knowledge graph types

LLMs for Knowledge Graph Construction and Reasoning: Recent Capabilities and Future Opportunities

May 22, 2023
YZ
Yuqi Zhu
🏛️ Zhejiang University | National University of Singapore

This work systematically evaluates large language models (LLMs) for knowledge graph (KG) construction and reasoning across four core tasks: entity/relation extraction, event identification, link prediction, and question answering. The study identifies a key insight: LLMs serve more effectively as *reasoning assistants* than few-shot extractors. To advance evaluation, the authors introduce the novel *virtual knowledge extraction* task and the VINE benchmark dataset. They further propose AutoKG—a multi-agent framework integrating prompt engineering, external knowledge retrieval, and dynamic verification—to enable end-to-end KG construction and reasoning. Experiments demonstrate that GPT-4–based AutoKG outperforms fine-tuned models on multiple reasoning benchmarks and achieves superior cross-domain generalization. The codebase and VINE dataset are publicly released to foster deeper integration of KGs and LLMs.

Few-shot LearningKnowledge Graph TasksLarge Language Models

Latest Papers

What's happening recently
View more

This work addresses the need for unified and efficient knowledge provisioning in large language models by proposing a novel architecture that integrates relational and property graph data models. The approach leverages record addresses from log files as immutable reference values in place of traditional foreign keys, enabling efficient graph-style link traversal instead of costly join queries while natively supporting triple-based knowledge representation. The resulting unified knowledge service framework combines the structural rigor of relational models with the flexible associative capabilities of graph models, significantly enhancing knowledge retrieval efficiency and effectively supporting knowledge integration and invocation in generative AI systems.

data integrationknowledge representationlarge language models

This work addresses the high integration and reuse costs, as well as limited adaptability, of current knowledge graph modeling practices—ranging from lightweight vocabularies to axiom-rich ontologies—particularly in the context of neuro-symbolic AI where flexibility to evolving requirements is crucial. To unify these diverse approaches, the paper proposes a novel “ontology continuum” framework, systematically characterizing existing modeling paradigms along two orthogonal dimensions: semantic–pragmatic and attributive–functional. The continuum is formally modeled using Formal Concept Analysis (FCA) and empirically validated through real-world engineering cases and provenance analysis. This study provides the first experience-driven conceptual foundation for knowledge graph re-engineering, identifies five key open challenges, and advances the synergistic application of knowledge graphs in neuro-symbolic and generative AI systems.

Data IntegrationGenAIKnowledge Graph

This work addresses the pervasive ambiguity of concepts in knowledge graphs arising from domain-specific interpretations, a challenge inadequately handled by existing approaches that treat domain as external metadata and thus fail to enforce assertion validity at the representation level. The paper proposes the Domain-Contextualized Concept Graph (DCG) framework, which uniquely integrates domain as an intrinsic structural component of knowledge representation. By leveraging Kripke-style semantics and a compact predicate system, DCG embeds domain context directly into relational representations, enabling truth evaluation, inference, and conflict detection within specific domains. The framework supports concept disambiguation, domain-specific assertion validation, and explicit cross-domain relation linking, while providing formal mappings to RDF, OWL, and relational databases. This yields a computable, structurally grounded mechanism for domain-aware constraints, effectively mitigating knowledge system failures caused by neglecting domain context.

concept ambiguitycontextual semanticsdomain-constrained knowledge representation

This work addresses the limitations of large language models in complex multi-hop knowledge graph reasoning—namely, insufficient flexibility, fragmented reasoning processes, and loss of intermediate information—by introducing the first end-to-end framework that internalizes multi-step graph reasoning as a unified “thinking” process within the model. Trained via reinforcement learning, the model dynamically explores and backtracks along reasoning paths, overcoming the constraints of conventional pipeline-based approaches. By seamlessly integrating large language models, knowledge graphs, and reinforcement learning, the proposed method achieves state-of-the-art or competitive performance across eight knowledge-intensive multi-hop reasoning benchmarks, substantially enhancing both the coherence and accuracy of the reasoning process.

intermediate information lossKnowledge Base Question AnsweringKnowledge Graphs

Existing knowledge graph benchmark datasets commonly lack complete ontological schema information, limiting their utility for evaluating algorithms that rely on semantic constraints or neuro-symbolic reasoning. To address this gap, this work proposes a workflow that jointly extracts both schema and factual triples from knowledge graphs to construct consistency-aware datasets. By leveraging the OWL ontology language and description logic-based reasoning mechanisms, the approach resolves inconsistencies and infers implicit knowledge. The project delivers the first systematically constructed, high-expressivity dataset that integrates a complete ontological schema with factual assertions, while also enriching existing benchmarks with schema information. All released resources support both logical reasoning services and tensor-based loading in mainstream machine learning frameworks, substantially enhancing the fidelity and comprehensiveness of algorithm evaluation.

datasetknowledge graphmachine learning

Hot Scholars

YC

Yupeng Cao

Stevens Institute of Technology
Natural Language ProcessingMultiModalTrustworthy AI
MS

Maosong Sun

Professor of Computer Science and Technology, Tsinghua University
Natural Language ProcessingArtificial IntelligenceSocial Computing
YY

Yukun Yan

Tsinghua University
Large Language Model
ZL

Zhenghao Liu

Northeastern University
NLPInformation Retrieval
WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc