MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion

πŸ“… 2026-07-16
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing diffusion-based approaches for multimodal knowledge graph completion directly denoise raw features, struggling to simultaneously address relation-dependent clue selection, cross-modal semantic alignment, and structure-aware generation, which often leads to semantic inconsistency and suboptimal performance. To overcome these limitations, this work proposes a novel β€œalign-then-diffuse” paradigm. It first employs a Relation-Adaptive Semantic Routing Mixture-of-Experts (RASR-MoE) to select relevant modal pathways and leverages a frozen multimodal large language model (MLLM) to achieve cross-modal semantic alignment. Subsequently, diffusion-based generation conditioned on the knowledge graph structure (KGDT) is performed in a unified latent space. By decoupling semantic integration from generation, the proposed method significantly outperforms strong baselines across three benchmark datasets, effectively enhancing both generative consistency and completion accuracy.
πŸ“ Abstract
Multimodal Knowledge Graph Completion (MKGC) requires inferring missing entities from structural, textual, and visual cues. Existing diffusion-based MKGC methods usually denoise directly on raw multimodal features. Such a design forces the denoiser to simultaneously perform relation-dependent cue selection, cross-modal semantic alignment, and structure-aware entity generation, which introduces noisy and semantically inconsistent conditions for diffusion and consequently leads to suboptimal completion performance. To address this limitation, we propose MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts (MGDT), a novel MKGC framework built on an align-then-diffuse paradigm. MGDT first employs a Relation-Adaptive Semantic Routing Mixture-of-Experts (RASR-MoE) module to select relation-relevant multimodal semantic transformation paths and suppress irrelevant modality interference. MGDT then uses a frozen Multimodal Large Language Model (MLLM) as a semantic anchor to align the routed multimodal representations into a unified latent space and reduce cross-modal semantic heterogeneity. Finally, a Knowledge Graph Diffusion Transformer (KGDT) performs graph-conditioned denoising generation in the aligned space to produce the missing entity representation. Experiments on three benchmark datasets show that MGDT consistently outperforms strong baselines.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Knowledge Graph Completion
Diffusion Model
Cross-modal Semantic Alignment
Relation-dependent Cue Selection
Semantic Heterogeneity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Relation-Adaptive Mixture-of-Experts
MLLM-Guided Alignment
Align-then-Diffuse Paradigm
Multimodal Knowledge Graph Completion
Diffusion Transformer
Xu Hou
Xu Hou
Professor of Xiamen University
bio-inspired design of advanced materialschemical modificationbiomedical engineeringmicrofluidicsmembrane science
M
Meiyu Liang
School of Computer Science, Beijing University of Posts and Telecommunications
Wei Huang
Wei Huang
UPS XI-INRS-ZJU
OptoelectronicsPulsed laser
Yawen Li
Yawen Li
Lawrence Technological University
Biomaterialstissue engineeringBioMEMS
Z
Zhe Xue
School of Computer Science, Beijing University of Posts and Telecommunications
W
Wu Liu
University of Science and Technology of China
G
Guanhua Ye
School of Computer Science (National Pilot Software Engineering School), Beijing University of Posts and Telecommunications
L
Lei Shi
Communication University of China
K
Kangkang Lu
Beijing University of Posts and Telecommunications