MSGAT: Multi-Head Spiking Graph Attention with Similarity-Space Fusion for Image-Text Retrieval

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of spiking neural networks (SNNs) in modeling semantic structures and fusing heterogeneous features for image-text retrieval by proposing the MSGAT framework. Methodologically, it introduces a dynamic multi-head spiking graph attention mechanism to enhance graph-structured representation learning. Additionally, a Sim-Fuse strategy is designed to perform event-driven fusion of multi-granularity features within the similarity space, effectively circumventing conflicts arising from directly merging heterogeneous representations. Experimental results demonstrate that the proposed approach surpasses existing SNN baselines and artificial neural network models of comparable scale on both the Flickr30K and MSCOCO datasets. Notably, it achieves a 55% reduction in energy consumption, enabling high-accuracy, energy-efficient cross-modal alignment.
📝 Abstract
Spiking neural networks (SNNs) offer an energy-efficient computing paradigm through sparse event-driven computation, showing great potential for efficient multimodal learning. However, applying SNNs to high-level multimodal tasks, such as image-text retrieval (ITR), remains challenging, since sparse spike representations make it difficult to capture semantic structures required for cross-modal alignment. Existing spiking ITR methods rely on local alignment and additional soft-label supervision during training, while lacking awareness of structural and multi-granularity relationships. To address these issues, we propose a Multi-head Spiking Graph Attention Network (\textbf{MSGAT}) for structural modeling and equip it with dynamic attention heads to capture complementary relational patterns and enable spike-driven graph reasoning and aggregation. However, within a two-branch multi-granularity fusion framework, the fine-grained spike representations generated by MSGAT are sparse and discrete, whereas the global representations are continuous, making conventional feature-level fusion susceptible to interference across heterogeneous representations. Therefore, we introduce \textbf{Sim-Fuse}, a similarity-space fusion alignment strategy integrating coarse- and fine-grained matching relations while avoiding direct fusion of heterogeneous representations. Experiments on Flickr30K and MSCOCO show our method outperforms ANN methods under matched settings and existing SNN retrieval baselines. Moreover, with only two time steps, our SNN achieves comparable or superior performance to its ANN counterpart while reducing theoretical module-level energy by 55\%. The code is provided in the Supplementary Materials.
Problem

Research questions and friction points this paper is trying to address.

Spiking Neural Networks
Image-Text Retrieval
Cross-modal Alignment
Multi-granularity Fusion
Heterogeneous Representations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spiking Neural Networks
Graph Attention
Image-Text Retrieval
Similarity-Space Fusion
Energy Efficiency
🔎 Similar Papers
No similar papers found.