Score
Designs and builds knowledge-graph schemas and instantiated graphs that integrate entities, attributes, and typed relations extracted from multiple modalities (for example text and images) by generating nodes, relation edges, and summarizing document- or scene-level structure. Implements cross-modal mention alignment, entity linking and canonicalization, relation extraction across modalities, and neuro-symbolic integration that combines learned representations with symbolic KG construction and reasoning.
This work proposes a novel knowledge graph representation paradigm that addresses the limitations of traditional approaches, which rely on predefined symbolic relations and consequently fail to capture the context-dependency, fine-grained semantics, and inherent uncertainty of real-world relationships—often leading to critical information loss. By reformulating relations as natural language descriptions rather than discrete symbols, the proposed framework leverages the generative and reasoning capabilities of large language models. It integrates prompt engineering with a minimal structural backbone, thereby harmonizing structured scaffolding with unstructured semantic expression. This hybrid design substantially enhances the semantic richness and context-awareness of relation modeling, offering a more adaptive and expressive pathway for knowledge graph construction in the era of large language models.
Existing knowledge graph embedding (KGE) methods struggle to effectively model multimodal entities and often process different modalities in isolation, resulting in weak cross-modal alignment and overly simplified semantic assumptions. This work proposes the first end-to-end joint representation learning framework that integrates vision-language models (VLMs) into multimodal KGE, leveraging VLMs’ inherent cross-modal alignment capabilities together with relational structure modeling from knowledge graphs to overcome modality isolation. Experimental results on WN9-IMG and two newly constructed art-domain multimodal knowledge graphs demonstrate that the proposed approach significantly outperforms current unimodal and multimodal KGE methods on link prediction tasks.
To address the scarcity of multimodal knowledge graphs (MMKGs) and the low quality of image–entity associations, this paper proposes an automatic framework for constructing high-quality MMKGs from unimodal knowledge graphs. The core innovation is the Visualized Structured Neighbor Selection (VSNS) method, which—novelty—decouples visuality-aware relation filtering (VNS) from structural neighborhood selection (SNS), enabling knowledge-context-driven image selection and generation. VSNS integrates knowledge graph embedding, relation-level visuality discrimination, neighborhood importance scoring, and multimodal prompt engineering to guide diffusion-based image generation. Experiments on MKG-Y and DB15K demonstrate that VSNS significantly improves both semantic relevance and structural consistency of generated images. Quantitative evaluations—including image–entity alignment metrics—and qualitative analyses consistently outperform existing baselines, validating the effectiveness of our approach in bridging unimodal knowledge graphs with rich, contextually grounded visual content.
This work addresses the limitations of traditional knowledge graph construction approaches, wherein structural decisions are hard-coded into rigid pipelines, resulting in tight coupling between schema and construction process and hindering support for ontology-level tasks. To overcome this, the authors propose an ontology-oriented construction framework featuring a novel intrinsic-relational routing mechanism. This mechanism dynamically assigns attributes to corresponding schema modules through iterative attribute classification, enabling a declarative and reusable decoupled design. The pipeline integrates rule-based cleaning, tool-augmented large language model–assisted annotation, and human review. Evaluated on Wikidata (January 2026), the resulting graph comprises 34 million nodes and 61.2 million edges, achieving 93.3% schema coverage and 98.0% module assignment accuracy, effectively supporting five ontology-level applications.
Existing large language models (LLMs) for knowledge graph construction suffer from a local text-centric perspective, hindering cross-document information integration and implicit relation discovery. To address this, we propose Graphusion: a zero-shot, end-to-end framework that leverages seed entities to guide LLM-based triple extraction and introduces a novel global fusion module—supporting entity consolidation, conflict resolution, and implicit relation discovery—thereby transcending conventional extraction paradigms. We formally define the zero-shot knowledge graph completion (KGC) task without fine-tuning and present TutorQA, the first expert-validated QA benchmark for education-domain KGC. Experiments show Graphusion achieves triple extraction scores of 2.92/3 (entities) and 2.37/3 (relations); improves subgraph completion accuracy by 9.2%; and significantly enhances downstream question answering performance.
Traditional knowledge graph construction methods struggle to balance structural consistency and contextual richness, as they are either constrained by the high cost of predefined ontologies or suffer from fragmentation due to schema-free extraction. This work proposes TRACE-KG, a novel framework that, for the first time, jointly generates a knowledge graph and a data-driven semantic schema without requiring any pre-specified ontology. The approach leverages multimodal joint modeling and structured qualifiers to represent conditional relationships, enabling end-to-end traceable knowledge extraction. Experimental results on complex technical documents demonstrate that the generated graphs significantly outperform existing ontology-driven and schema-free methods in terms of structural coherence, semantic richness, and traceability back to the source documents.
This study addresses the limitations of existing knowledge graph embedding methods, which primarily model at the triple level and struggle to capture graph-level semantic similarity, while traditional structural alignment approaches often lack semantic consistency. The work presents the first systematic evaluation of knowledge graph embeddings for graph-to-graph semantic matching tasks, introducing a novel dataset constructed via document rewriting to reflect semantic similarity. Two lightweight and efficient embedding-based scoring functions are proposed: Maximum Pairwise Entity Similarity (EmbPairSim) and Frequency-Weighted Centroid Similarity (AvgEmbSim). Experimental results on WikiText-2 and CC-News demonstrate that EmbPairSim achieves up to a 5.3-point MRR improvement over Sentence-BERT, confirming the effectiveness of knowledge graph embeddings as compact yet informative signals for graph-level semantic matching.
This work addresses the “hairball problem” in large-scale knowledge graphs, where dense nodes and relations obscure semantic information and hinder interpretability. To tackle this challenge, the authors propose an interactive semantic aggregation method powered by large language models (LLMs), enabling user-driven compression of graph structures. The approach dynamically aggregates local subgraphs into interpretable super-nodes and super-edges while preserving traceability to original triples and source documents. By integrating LLMs, knowledge graph construction, and interactive visualization, the method effectively surfaces high-level insights from document collections in applications such as movie review analysis and intelligence assessment, significantly enhancing both the comprehensibility and explainability of complex knowledge graphs.
This work addresses the challenges of multimodal knowledge graph completion under relation sparsity, where conventional embedding methods suffer performance degradation and large language models lose critical topological and visual information through linearization. To overcome these limitations, the paper proposes ViSR-KGC, which uniquely visualizes query-relevant subgraphs as images and constructs a unified prompt by integrating entity images, textual descriptions, and candidate answers to guide reasoning with a vision-language model. By synergistically combining multimodal embedding learning, subgraph extraction, graph layout visualization, and prompt engineering, ViSR-KGC effectively leverages global topological structure, local multimodal evidence, and pretrained commonsense knowledge. The method significantly outperforms existing embedding-based and large language model approaches in sparse relational scenarios.