When Graph meets Multimodal: Benchmarking and Meditating on Multimodal Attributed Graphs Learning

📅 2024-10-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The absence of standardized benchmarks and systematic evaluation protocols for multimodal attributed graphs (MAGs) hinders progress in multimodal graph learning. Method: We introduce MAGB—the first open-source, multi-domain MAG benchmark—and propose a unified evaluation framework integrating graph neural networks (GNNs), vision-language models (VLMs), multimodal embedding alignment, and zero-shot inference. We systematically compare two paradigms: GNN-as-Predictor and VLM-as-Predictor. Contributions/Results: Key findings include: (1) modality importance is domain-dependent; (2) multimodal embeddings elevate GNN performance ceilings but introduce modality bias; (3) VLMs effectively mitigate image–text imbalance. Experiments show that joint image–text modeling improves average node classification accuracy by 12.7%; VLMs significantly enhance generalization under low-resource conditions. This work establishes the first comprehensive MAG evaluation suite, advancing multimodal graph learning research in domains such as social networks and e-commerce.

Technology Category

Machine Learning: Multimodal LearningComputer Vision: Multi-modal VisionData Mining & Knowledge Management: Mining of Visual, Multimedia & Multimodal Data

Application Category

Graph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web data
📝 Abstract
Multimodal Attributed Graphs (MAGs) are ubiquitous in real-world applications, encompassing extensive knowledge through multimodal attributes attached to nodes (e.g., texts and images) and topological structure representing node interactions. Despite its potential to advance diverse research fields like social networks and e-commerce, MAG representation learning (MAGRL) remains underexplored due to the lack of standardized datasets and evaluation frameworks. In this paper, we first propose MAGB, a comprehensive MAG benchmark dataset, featuring curated graphs from various domains with both textual and visual attributes. Based on MAGB dataset, we further systematically evaluate two mainstream MAGRL paradigms: $ extit{GNN-as-Predictor}$, which integrates multimodal attributes via Graph Neural Networks (GNNs), and $ extit{VLM-as-Predictor}$, which harnesses Vision Language Models (VLMs) for zero-shot reasoning. Extensive experiments on MAGB reveal following critical insights: $ extit{(i)}$ Modality significances fluctuate drastically with specific domain characteristics. $ extit{(ii)}$ Multimodal embeddings can elevate the performance ceiling of GNNs. However, intrinsic biases among modalities may impede effective training, particularly in low-data scenarios. $ extit{(iii)}$ VLMs are highly effective at generating multimodal embeddings that alleviate the imbalance between textual and visual attributes. These discoveries, which illuminate the synergy between multimodal attributes and graph topologies, contribute to reliable benchmarks, paving the way for future MAG research. The MAGB dataset and evaluation pipeline are publicly available at https://github.com/sktsherlock/MAGB.
Problem

Research questions and friction points this paper is trying to address.

Benchmarking Multimodal Attributed Graphs Learning
Evaluating GNN and VLM paradigms
Addressing modality biases in graph learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proposes MAGB benchmark for multimodal graphs.
Evaluates GNN-as-Predictor and VLM-as-Predictor paradigms.
Discovers VLMs enhance multimodal embedding effectiveness.
💼 Related Jobs
No related jobs found.
Central South University | Microsoft Research Asia | Microsoft
H
Hao Yan
Central South University
Chaozhuo Li
Chaozhuo Li
Microsoft Research Aisa
J
Jun Yin
Central South University
Z
Zhigang Yu
Central South University
W
Weihao Han
Microsoft
M
Mingzheng Li
Microsoft
Z
Zhengxin Zeng
Microsoft
H
Hao Sun
Microsoft
S
Senzhang Wang
Central South University