Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data

๐Ÿ“… 2026-10-07
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenge of unsupervised cross-modal alignment between independently trained models, which typically requires extensive paired data. Grounded in the Platonic Representation Hypothesis, this work proposes a pairing-free alignment method that establishes cross-modal correspondences by exploiting geometric similarities within representation spaces. Specifically, orthogonal mappings are estimated via the Wasserstein Procrustes algorithm combined with geometric initialization techniques to effectively align embedding spaces. Notably, this research provides the first demonstration that coarse-grained cross-modal alignment can be achieved using geometric initialization alone. Extensive experiments across multiple datasets and modalities reveal that the proposed approach significantly outperforms existing baselines under extremely limited paired-data scenarios. Furthermore, its practical utility is validated through successful application to text-to-image generation tasks.
๐Ÿ“ Abstract
Multimodal representations enable zero-shot classification and retrieval, but aligning independently trained models usually requires large amounts of paired data. Yet, the Platonic Representation Hypothesis suggests that models trained on different modalities may converge spontaneously toward a shared representation geometry. But then, do we even need paired examples for cross-modal alignment? Remarkably, we show that paired examples are unnecessary for coarse cross-modal alignment. Our simple Wasserstein Procrustes method with a coarse geometric initialization aligns two disjoint embedding sets by estimating a single orthogonal map without seeing any pairs. Across datasets, modalities, and unimodal models, we show that we can consistently align independently trained representations without pairs, and standard geometric alignment metrics accurately predict when this is possible. Nevertheless, we can naturally benefit from paired examples. In the very few-pair regime, our method substantially outperforms existing ones, while staying competitive with pair-based methods with more added examples. Finally, we demonstrate that the resulting alignments can enable text-to-image generation without paired examples. These results show that independently trained models often share enough geometry to establish cross-modal correspondence with little or no paired data.
Problem

Research questions and friction points this paper is trying to address.

cross-modal alignment
paired data
multimodal representations
representation geometry
zero-shot
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Modal Alignment
Wasserstein Procrustes
Platonic Representation Hypothesis
Unpaired Data
Shared Geometry
๐Ÿ”Ž Similar Papers
No similar papers found.