Score
Designs, builds, and operates structured knowledge artifacts (knowledge bases and knowledge graphs) by defining schemas/ontologies and representation formats, implementing entity and relation extraction, canonicalization, and linking, and creating ingestion, transformation, storage, indexing, and query/retrieval pipelines. It also covers integrating and harmonizing data across sources, managing and maintaining versions and updates, and providing tools for knowledge capture, management, and lifecycle engineering.
This paper addresses the limitations of traditional rule- and statistics-based approaches in knowledge graph (KG) construction—namely, ontology engineering, knowledge extraction, and knowledge fusion. It proposes a novel “language-driven generative KG construction” paradigm that unifies schema-based and schema-agnostic methods, establishing synergistic mechanisms between large language models (LLMs) and symbolic KGs for structured organization and open-ended semantic expression. The method integrates prompt engineering, knowledge representation learning, automated reasoning, and multimodal modeling to enable dynamic, interpretable knowledge acquisition and fusion. The study comprehensively surveys technical pathways and bottlenecks, identifying three key research directions: LLM reasoning enhancement via KGs, agent memory modeling with KGs, and multimodal KG construction. Ultimately, this work advances the development of adaptive, neuro-symbolic intelligent knowledge systems.
In the era of large language models, traditional record-centric data engineering struggles to meet the demand for organizational knowledge as executable infrastructure. This work proposes a novel paradigm—knowledge architecture—that systematically reimagines core data engineering mechanisms by upgrading ETL, data lineage, and catalogs into knowledge ingestion, change detection, provenance, and knowledge catalogs. It introduces knowledge views and a three-tier layered model (raw–refined–operational) to structure knowledge effectively. By integrating emerging standards such as LLM Wiki and Open Knowledge Format (OKF), this study formally defines knowledge architecture for the first time and establishes a theoretical framework that supports knowledge representation, governance, and operational delivery, enabling direct invocation of organizational knowledge by humans, agents, workflows, and models alike.
This work addresses the limitations of traditional knowledge graph construction approaches, wherein structural decisions are hard-coded into rigid pipelines, resulting in tight coupling between schema and construction process and hindering support for ontology-level tasks. To overcome this, the authors propose an ontology-oriented construction framework featuring a novel intrinsic-relational routing mechanism. This mechanism dynamically assigns attributes to corresponding schema modules through iterative attribute classification, enabling a declarative and reusable decoupled design. The pipeline integrates rule-based cleaning, tool-augmented large language model–assisted annotation, and human review. Evaluated on Wikidata (January 2026), the resulting graph comprises 34 million nodes and 61.2 million edges, achieving 93.3% schema coverage and 98.0% module assignment accuracy, effectively supporting five ontology-level applications.
Entity disambiguation and linking in IT domains suffer from poor domain adaptability and difficulty in incorporating domain-specific knowledge when relying solely on general-purpose knowledge graphs (e.g., Wikidata, DBpedia). Method: This paper proposes a lightweight, extensible ontology construction method tailored for the IT domain. Starting from general Linked Open Data (LOD) resources, it employs a domain-agnostic pipeline and—novelly—integrates an IT-specific terminology lexicon to drive ontology schema expansion. The approach synergistically combines SPARQL querying, RDF reasoning, and ontology alignment. Contribution/Results: The resulting paradigm balances generality and domain specificity, significantly improving accuracy in IT entity disambiguation and linking. It establishes a low-barrier, reusable ontology engineering framework that supports continuous injection of proprietary domain knowledge, thereby enabling sustainable, scalable domain ontology development.
Legacy systems written in COBOL, PL/I, or Assembly—common in banking and telecommunications—are often undocumented and lack original developers, hindering comprehension and modernization. Method: This paper proposes a multi-language, cross-platform, customizable framework for constructing software knowledge graphs and interactively defining architectural boundaries. It integrates static code analysis, data schema parsing, and custom ontology modeling to enable expert-guided, incremental analysis of source code and data architecture, automatically identifying business- and data-driven logical boundaries and visualizing cross-boundary dependencies. Contribution/Results: The framework introduces the first knowledge-graph-driven approach for progressive modernization path planning and impact analysis. Evaluated on two real-world industrial systems, it significantly improves system understanding efficiency and enhances the accuracy of modernization strategy design.
To address low automation, poor interpretability, and semantic incompatibility with Wikidata in knowledge graph (KG) construction, this paper proposes an ontology-driven large language model (LLM) approach. First, domain scope and relations are lightweightly extracted from Competency Questions (CQs). Second, extracted relations are bidirectionally mapped to the Wikidata ontology to achieve semantic alignment. Third, LLMs are guided—under ontology constraints—to generate structured subject-predicate-object triples. The key contribution is the first CQ-driven ontology construction paradigm, uniquely balancing automation and interpretability. Experiments on standard benchmarks demonstrate state-of-the-art performance; the resulting KG exhibits high logical consistency and cross-system interoperability, significantly reducing reliance on manual annotation. The method enables a scalable, reusable KG construction pipeline.
This work addresses the limitations of existing ontology documentation tools in supporting modular modeling and human readability, particularly in handling cross-module entities and annotations. To overcome these challenges, the authors refactor and extend the LODE framework by introducing a modular Reader-Model-Viewer architecture that decouples parsing, modeling, and rendering components. Implemented as a web service, the new framework provides enhanced capabilities for generating OWL ontology documentation, featuring dedicated entity pages, RDF provenance tracking, and Markdown-based rendering. These improvements significantly increase the intelligibility and reusability of modular scientific knowledge graph ontologies. The framework has been successfully applied to the documentation of the SKG-O ontology, demonstrating its practical utility and effectiveness.
This work addresses the challenge of architectural knowledge fragmentation across heterogeneous software artifacts, which often leads to architectural degradation due to inconsistent evolution. The paper presents the first end-to-end automated framework for architectural knowledge management, featuring specialized extractors that harvest knowledge from diverse sources, a unified representation model, and integrated mechanisms for consistency checking and repair to preserve knowledge integrity. Innovatively leveraging Retrieval-Augmented Generation (RAG), the framework enables compliance verification, change impact analysis, and natural language question answering grounded in a structured knowledge base. An experimental prototype demonstrates the approach’s effectiveness in enhancing architectural understanding, maintenance efficiency, and intelligent interaction.
This work addresses the need for unified and efficient knowledge provisioning in large language models by proposing a novel architecture that integrates relational and property graph data models. The approach leverages record addresses from log files as immutable reference values in place of traditional foreign keys, enabling efficient graph-style link traversal instead of costly join queries while natively supporting triple-based knowledge representation. The resulting unified knowledge service framework combines the structural rigor of relational models with the flexible associative capabilities of graph models, significantly enhancing knowledge retrieval efficiency and effectively supporting knowledge integration and invocation in generative AI systems.
This study addresses persistent challenges in scientific communication and aerospace engineering—namely data silos, insufficient collaboration incentives, and legal barriers—that hinder the implementation of FAIR (Findable, Accessible, Interoperable, Reusable) principles. To overcome these limitations, this work proposes a novel, scalable knowledge infrastructure framework that integrates human–AI collaboration, knowledge graphs, and user-centered design across technological, social, and legal dimensions. The framework encompasses automated information processing workflows, a wiki-style digital library, and demand-driven interactive interfaces. Pilot implementations demonstrate its effectiveness in consolidating fragmented knowledge resources and establishing a viable collaborative paradigm for sparsely networked domains. Nevertheless, institutional and sociocultural barriers remain significant and require further intervention to fully realize the framework’s potential.
This study addresses the lack of systematic guidance on contextualization strategies for large language model (LLM) agents operating in structured data environments, particularly concerning effectiveness and efficiency across multi-file, large-scale schemas. Using SQL generation as a proxy task, the work presents the first systematic evaluation of eleven models across four context formats—YAML, Markdown, JSON, and TOON—at schema scales ranging from 10 to 10,000 tables. The findings reveal that model capability tiers critically determine optimal context architecture: tailored strategies significantly improve performance, with state-of-the-art models gaining 2.7% accuracy under native file-based contexts, while open-source models average a 7.7% decline. Moreover, native file-based agents scale efficiently to ten-thousand-table schemas while maintaining high navigation accuracy.