Score
Designs and implements graph-structured representation learning systems that jointly encode multiple data modalities (e.g., text, images) and temporal interaction signals into node and edge embeddings, including neighborhood-interaction and weighted-process graph constructions. Builds or analyzes models that use these multimodal graph representations to forecast node- or edge-level outcomes such as future content popularity or other graph-based predictions.
The absence of standardized benchmarks and systematic evaluation protocols for multimodal attributed graphs (MAGs) hinders progress in multimodal graph learning. Method: We introduce MAGB—the first open-source, multi-domain MAG benchmark—and propose a unified evaluation framework integrating graph neural networks (GNNs), vision-language models (VLMs), multimodal embedding alignment, and zero-shot inference. We systematically compare two paradigms: GNN-as-Predictor and VLM-as-Predictor. Contributions/Results: Key findings include: (1) modality importance is domain-dependent; (2) multimodal embeddings elevate GNN performance ceilings but introduce modality bias; (3) VLMs effectively mitigate image–text imbalance. Experiments show that joint image–text modeling improves average node classification accuracy by 12.7%; VLMs significantly enhance generalization under low-resource conditions. This work establishes the first comprehensive MAG evaluation suite, advancing multimodal graph learning research in domains such as social networks and e-commerce.
Real-world multimodal data exhibit complex cross-modal structural relationships—such as coreference and contextual dependencies—that extend far beyond simple image–text alignment. Existing methods treat graphs as isolated modalities and neglect multi-hop neighbor interactions, leading to fragmented semantic understanding. To address this, we propose a unified framework that jointly optimizes multi-hop graph structural modeling and multimodal fusion. Specifically, we introduce Hop-Diffused Attention to explicitly distinguish multi-hop neighbors and design MM-QFormer for principled, modality-aware fusion. Our architecture integrates graph neural networks, causal masking, diffusion-based mechanisms, and multi-mapping query Transformers. Evaluated on both generative and discriminative multimodal tasks, our model—despite its smaller scale—outperforms state-of-the-art vision-language models and multimodal graph models, achieving an average performance gain of 6.93%.
This work addresses the limitations of existing methods in handling heterogeneous node features—such as images and text—in multimodal graphs, where inflexible and inefficient intra- and inter-modal message passing hinders performance. To overcome this, we propose the Dynamic Information Pathway (DiP) framework, which introduces modality-specific pseudo-nodes to construct dynamic, sparse information pathways within a shared state space. This design enables adaptive intra-modal message routing and efficient inter-modal dependency modeling. Notably, DiP achieves cross-modal adaptive propagation with linear complexity, circumventing the constraints of static architectures or dense attention mechanisms. Extensive experiments demonstrate that DiP significantly outperforms state-of-the-art approaches on multiple benchmark datasets for both link prediction and node classification tasks.
Systematic integration of visual information with graph structure remains underexplored in graph machine learning. Method: We introduce MM-GRAPH, the first benchmark for vision–text–structure multimodal graph learning, comprising seven real-world datasets and supporting standardized evaluation across node classification, link prediction, and other core tasks. It uniquely incorporates high-dimensional visual features as node attributes and provides an end-to-end, reproducible evaluation framework featuring GNN backbones, cross-modal alignment and fusion modules, and unified preprocessing pipelines. Contribution/Results: Empirical analysis demonstrates that incorporating visual modality substantially improves model generalization—particularly under sparse labeling and long-tailed class distributions. MM-GRAPH fills a critical gap in multimodal graph learning benchmarks and establishes a rigorous foundation for both methodological development and diagnostic analysis.
This work addresses the limitation of existing graph learning approaches, which typically operate in isolation within a single modality and task, thereby hindering the cross-task and cross-modal reuse of structural knowledge. To overcome this, the authors propose G-Substrate, a novel framework that models graph structures as persistent, shareable substrates. By unifying structural patterns and employing a role-interleaved training strategy, G-Substrate enables collaborative learning across multiple tasks and modalities. This approach facilitates the continuous accumulation and transfer of graph-structured knowledge, consistently outperforming both isolated training and conventional multi-task learning methods across diverse domains, modalities, and tasks.
This work addresses the challenge in multimodal attributed graphs where structural-induced semantics and modality-intrinsic semantics contribute differently to downstream tasks, yet conventional coupled approaches fail to disentangle them, limiting cross-modal fusion effectiveness. To this end, we propose the first pretraining framework based on graph spectral decomposition: leveraging scalable Chebyshev filters to decompose node signals of each modality into frequency bands, thereby constructing band-resolved modality tokens. We incorporate graph spectral priors to design a frequency-routing mechanism that promotes structural consistency while preserving modality specificity. Furthermore, a topology-conditioned routing strategy evaluates coupling reliability, enabling differentiated interactions across frequency bands and modalities. Extensive experiments demonstrate that our method achieves state-of-the-art performance on multiple multimodal attributed graph benchmarks, significantly improving both graph-level and modality-level task outcomes.
Existing deep learning models struggle to effectively encode spatial, topological, and semantic structural information inherent in images. This work systematically evaluates the impact of various visual graph construction strategies on image classification performance within a unified three-layer Graph Convolutional Network (GCN) framework. For the first time, it demonstrates that the graph structure itself plays a decisive role in model performance. The study underscores the critical importance of the graph construction preprocessing stage, providing empirical evidence that well-designed graph structures substantially enhance classification accuracy. These findings offer both methodological guidance and practical justification for graph structure selection and preprocessing in visual graph neural networks.
This work addresses the limitation of existing methods in multimodal attributed graphs, which employ a uniform message-passing mechanism that fails to differentiate the heterogeneous influence of neighbors across modalities, thereby blurring modality-specific signals. To overcome this, the authors propose RoleMAG, a novel framework that introduces a neighbor role-aware mechanism to dynamically categorize neighbors into three types—shared, complementary, or heterogeneous—and routes them to corresponding modality-specific propagation channels. This enables fine-grained, modality-aware information aggregation, effectively supporting cross-modal completion while preventing heterogeneous neighbors from disrupting the shared smoothness assumption. Extensive experiments demonstrate that RoleMAG achieves state-of-the-art performance on RedditS and Bili_Dance datasets and remains competitive on Toys. Ablation studies and robustness analyses further validate the effectiveness of the proposed approach.