🤖 AI Summary
To address the limitation of conventional distance metrics in distinguishing semantically dissimilar yet visually similar samples in multimodal learning, this paper introduces geodesic distance—previously unexplored in multimodal similarity modeling—to explicitly capture semantic discrepancies embedded in nonlinear manifold structures. Methodologically, we construct a hierarchical k-nearest-neighbor graph and employ efficient shortest-path algorithms (Dijkstra/Floyd) to compute geodesic distances. Furthermore, we propose a dynamic graph incremental update mechanism that preserves geometric fidelity while significantly improving computational efficiency. Extensive experiments on cross-modal retrieval and alignment tasks demonstrate that our approach consistently outperforms state-of-the-art baselines, validating the effectiveness, robustness, and generalizability of geodesic distance for modeling complex semantic relationships in multimodal representation learning.
📝 Abstract
Geodesic distance serves as a reliable means of measuring distance in nonlinear spaces, and such nonlinear manifolds are prevalent in the current multimodal learning. In these scenarios, some samples may exhibit high similarity, yet they convey different semantics, making traditional distance metrics inadequate for distinguishing between positive and negative samples. This paper introduces geodesic distance as a novel distance metric in multi-modal learning for the first time, to mine correlations between samples, aiming to address the limitations of common distance metric. Our approach incorporates a comprehensive series of strategies to adapt geodesic distance for the current multimodal learning. Specifically, we construct a graph structure to represent the adjacency relationships among samples by thresholding distances between them and then apply the shortest-path algorithm to obtain geodesic distance within this graph. To facilitate efficient computation, we further propose a hierarchical graph structure through clustering and combined with incremental update strategies for dynamic status updates. Extensive experiments across various downstream tasks validate the effectiveness of our proposed method, demonstrating its capability to capture complex relationships between samples and improve the performance of multimodal learning models.