Score
Design, build, and analyze components that connect applications or models to external structured knowledge resources (ontologies, knowledge graphs, or databases) by mapping schemas and entity identifiers, implementing entity linking and canonicalization, translating and routing queries, and supporting retrieval, updates, and consistency checks; evaluate coverage, provenance, and the effect of grounded knowledge on downstream system behavior.
This study addresses key challenges in deeply integrating large language models (LLMs) with structured knowledge systems—particularly knowledge graphs—including knowledge accuracy, dynamic updating, trustworthy reasoning, and ethical governance. Methodologically, it introduces the first multidimensional evaluation framework for LLM–knowledge base integration, formalizing three core benefits: data contextualization, precision enhancement, and knowledge utilization efficiency, while identifying critical gaps in scalability, real-time knowledge updating, and neuro-symbolic synergy. The approach unifies knowledge graph embedding, retrieval-augmented generation (RAG), prompt engineering, knowledge distillation, and explainability analysis to balance logical rigor with generative flexibility. Drawing on a systematic review of 200+ scholarly works, the study establishes a taxonomy and derives six actionable, industry-ready implementation guidelines. Results provide reusable integration paradigms and risk-mitigation pathways for high-stakes domains including finance, healthcare, and public administration.
This work addresses the limitations of traditional knowledge graph construction approaches, wherein structural decisions are hard-coded into rigid pipelines, resulting in tight coupling between schema and construction process and hindering support for ontology-level tasks. To overcome this, the authors propose an ontology-oriented construction framework featuring a novel intrinsic-relational routing mechanism. This mechanism dynamically assigns attributes to corresponding schema modules through iterative attribute classification, enabling a declarative and reusable decoupled design. The pipeline integrates rule-based cleaning, tool-augmented large language model–assisted annotation, and human review. Evaluated on Wikidata (January 2026), the resulting graph comprises 34 million nodes and 61.2 million edges, achieving 93.3% schema coverage and 98.0% module assignment accuracy, effectively supporting five ontology-level applications.
Legacy systems written in COBOL, PL/I, or Assembly—common in banking and telecommunications—are often undocumented and lack original developers, hindering comprehension and modernization. Method: This paper proposes a multi-language, cross-platform, customizable framework for constructing software knowledge graphs and interactively defining architectural boundaries. It integrates static code analysis, data schema parsing, and custom ontology modeling to enable expert-guided, incremental analysis of source code and data architecture, automatically identifying business- and data-driven logical boundaries and visualizing cross-boundary dependencies. Contribution/Results: The framework introduces the first knowledge-graph-driven approach for progressive modernization path planning and impact analysis. Evaluated on two real-world industrial systems, it significantly improves system understanding efficiency and enhances the accuracy of modernization strategy design.
This work addresses the challenge of architectural knowledge fragmentation across heterogeneous software artifacts, which often leads to architectural degradation due to inconsistent evolution. The paper presents the first end-to-end automated framework for architectural knowledge management, featuring specialized extractors that harvest knowledge from diverse sources, a unified representation model, and integrated mechanisms for consistency checking and repair to preserve knowledge integrity. Innovatively leveraging Retrieval-Augmented Generation (RAG), the framework enables compliance verification, change impact analysis, and natural language question answering grounded in a structured knowledge base. An experimental prototype demonstrates the approach’s effectiveness in enhancing architectural understanding, maintenance efficiency, and intelligent interaction.
This work addresses the challenge of integrating and retrieving multi-source heterogeneous data arising from schema inconsistencies by proposing an “executable schema contract” mechanism. This approach enables structure-aware automatic knowledge graph construction through a combination of closed-world field catalogs, deterministic structural analysis (e.g., primary/foreign key detection and source hierarchy identification), and monotonic extension protocols. It integrates large language model–constrained schema discovery, schema-guided information extraction and deduplication, and a multi-tool agent routing strategy that supports structured queries, graph traversal, and vector search. Evaluated on four question-answering benchmarks, the method achieves significantly superior zero-shot performance compared to pure retrieval and decomposition-based baselines. Ablation studies confirm that schema-conditioned routing, structural reasoning, and schema-guided construction are critical to its performance gains.
Entity disambiguation and linking in IT domains suffer from poor domain adaptability and difficulty in incorporating domain-specific knowledge when relying solely on general-purpose knowledge graphs (e.g., Wikidata, DBpedia). Method: This paper proposes a lightweight, extensible ontology construction method tailored for the IT domain. Starting from general Linked Open Data (LOD) resources, it employs a domain-agnostic pipeline and—novelly—integrates an IT-specific terminology lexicon to drive ontology schema expansion. The approach synergistically combines SPARQL querying, RDF reasoning, and ontology alignment. Contribution/Results: The resulting paradigm balances generality and domain specificity, significantly improving accuracy in IT entity disambiguation and linking. It establishes a low-barrier, reusable ontology engineering framework that supports continuous injection of proprietary domain knowledge, thereby enabling sustainable, scalable domain ontology development.
In the era of large language models, traditional record-centric data engineering struggles to meet the demand for organizational knowledge as executable infrastructure. This work proposes a novel paradigm—knowledge architecture—that systematically reimagines core data engineering mechanisms by upgrading ETL, data lineage, and catalogs into knowledge ingestion, change detection, provenance, and knowledge catalogs. It introduces knowledge views and a three-tier layered model (raw–refined–operational) to structure knowledge effectively. By integrating emerging standards such as LLM Wiki and Open Knowledge Format (OKF), this study formally defines knowledge architecture for the first time and establishes a theoretical framework that supports knowledge representation, governance, and operational delivery, enabling direct invocation of organizational knowledge by humans, agents, workflows, and models alike.
This study addresses the gap between existing Bodies of Knowledge (BoKs) in computing and their effective translation into assessable, competency-oriented curricula. The authors propose an innovative competency-mapping methodology that systematically aligns BoKs with a structured competency framework, resulting in a five-year engineering curriculum encompassing 23 core competencies, organized into five thematic modules and three specialization tracks, and integrating mandatory work-integrated learning projects. To support explicit linkage, collaborative maintenance, and continuous evolution of knowledge-to-competency mappings, the authors developed ISANUMpedia—a semantic web–based collaborative platform. Implemented in the ISANUM engineering degree program, this approach successfully mapped 494 knowledge topics to the 23 competencies, significantly enhancing the curriculum’s professional relevance and assessability.
This work addresses the limitations of existing ontology documentation tools in supporting modular modeling and human readability, particularly in handling cross-module entities and annotations. To overcome these challenges, the authors refactor and extend the LODE framework by introducing a modular Reader-Model-Viewer architecture that decouples parsing, modeling, and rendering components. Implemented as a web service, the new framework provides enhanced capabilities for generating OWL ontology documentation, featuring dedicated entity pages, RDF provenance tracking, and Markdown-based rendering. These improvements significantly increase the intelligibility and reusability of modular scientific knowledge graph ontologies. The framework has been successfully applied to the documentation of the SKG-O ontology, demonstrating its practical utility and effectiveness.
This study addresses the fragmentation of research software and its associated scholarly resources—such as publications and datasets—across disparate platforms, which hinders reproducibility and cross-domain analysis due to a lack of unified semantic links. To bridge this gap, the authors construct a large-scale RDF knowledge graph comprising 81 million triples, integrating approximately 200,000 GitHub repositories with external academic knowledge graphs including SemOpenAlex, LPWC, and MLSea-KG. This integration enables unified semantic modeling of software alongside scholarly entities such as authors, papers, and datasets. The resulting knowledge graph supports cross-platform provenance tracing and complex semantic queries, significantly enhancing the capacity to assess software reproducibility and analyze its long-term sustainability within the scientific ecosystem.
This study addresses the persistent challenge of efficiently translating biomedical knowledge into actionable outcomes, which is hindered by technical and organizational barriers in data integration. It introduces, for the first time, a systematic engineering paradigm centered on “knowledge assembly,” drawing on software engineering principles of composability and reproducibility, with an emphasis on the construction process rather than static artifacts. The proposed approach integrates key technologies—including identifier mapping, entity disambiguation, schema alignment, evidence provenance, typed namespaces, and standardized exchange formats—to design knowledge infrastructure that supports service composition and reproducible pipelines. Through an analysis of six representative knowledge graph systems, the work identifies eight open engineering challenges, offering both domain-specific guidance and a research agenda for advancing biomedical knowledge graph development through software engineering methodologies.