Score
Design and implement neural projection layers that map heterogeneous or modality-specific input representations into shared embedding spaces, choosing layer architectures, output dimensionality, normalization, and parameterization to preserve retrieval-relevant geometric relations. Specify and integrate training objectives and mechanisms (e.g., contrastive losses, regularization, alignment constraints) so the projections align modalities, maintain distances or neighborhood structure, and support downstream similarity or retrieval tasks.
Existing approaches typically rely solely on the final layer of deep models or employ simplistic fusion of shallow representations, overlooking the fact that task-relevant information is non-monotonically distributed across intermediate layers and cannot be reliably recovered through naive aggregation. This work systematically uncovers, for the first time, the distributional patterns of informative representations in intermediate layers and introduces a Layer-wise Optimal Embedding Selection (LOES) strategy grounded in spectral analysis, complemented by a geometric regularization loss (GeoReg). By leveraging constructive spectral methods along with orthogonality and isotropy constraints, the proposed framework precisely identifies task-discriminative subspaces and stabilizes their geometric structure. The method consistently outperforms baselines across diverse architectures, modalities, and data scales, exhibits performance gains with increasing model depth, and enables interpretable cross-lingual and cross-modal analyses.
Cross-modal retrieval suffers from semantic misalignment due to geometric inconsistency between image and text representations. This work formulates image–text matching as a metric alignment problem in the embedding space. We systematically demonstrate that: (1) the Wasserstein distance quantitatively characterizes inter-modal distributional discrepancy; (2) cosine similarity exhibits superior robustness over Euclidean distance, KL divergence, and other conventional metrics for alignment; and (3) standard MLPs inadequately capture complex cross-modal interactions. To address these issues, we propose a lightweight neural metric learning framework that integrates vision-language models with unimodal encoders, jointly optimizing four classical metrics and two learnable neural metrics. Our approach achieves significant improvements in Recall@K across standard cross-modal retrieval benchmarks. Moreover, it provides practical, deployment-oriented alignment evaluation criteria and architectural design guidelines.
Embedding models trained independently on similar data capture stable semantic meanings but yield inconsistent representation spaces, hindering interoperability across models. This work addresses compatibility challenges in multimodal search and model upgrades via orthogonal transformation-based embedding alignment. Theoretically, we derive the first tight Procrustes alignment error bound, proving the existence of a near-isometric orthogonal transformation that approximately preserves pairwise inner products—establishing rigorous theoretical foundations for alignment. Methodologically, we employ efficient Procrustes analysis as a post-hoc alignment procedure, preserving the intrinsic geometric structure of each embedding space while enabling cross-model alignment. Experiments demonstrate substantial improvements in model retraining compatibility, text retrieval fusion accuracy, and cross-modal search performance; our method achieves state-of-the-art results in hybrid multimodal search.
Multimodal contrastive learning, widely adopted for representation alignment, often fails to achieve semantic consistency because standard contrastive losses maximize mutual information without suppressing modality-specific information. Method: We introduce the information bottleneck principle into multimodal alignment for the first time, proposing a differentiable variational regularizer that explicitly enforces modality-invariant representations and suppresses modality-specific features within the contrastive learning framework. Contribution/Results: Our method requires no additional annotations and significantly improves alignment accuracy and semantic consistency in controlled ablation studies and cross-modal retrieval tasks. Empirical results demonstrate both the effectiveness and generalizability of information-bottleneck-driven regularization for multimodal representation learning.
Neural models trained on identical tasks exhibit geometrically heterogeneous latent representations due to stochasticity and architectural differences, rendering cross-model representations incomparable. To address this, we propose relative geodesic representations grounded in pullback metrics—a novel application of differential-geometric pullback metrics to latent space alignment—explicitly modeling intrinsic geometric transformations between latent manifolds of distinct models. Unlike conventional linear alignment methods, our approach operates without supervision and generalizes across diverse architectures and pretraining paradigms. Experiments on autoencoders and vision foundation discriminative models demonstrate substantial improvements in cross-model retrieval accuracy and model stitching performance. Moreover, the method scales effectively to large-scale settings.
This work addresses the degradation of cross-modal alignment in multimodal embedding spaces caused by modality bias and noisy supervision signals. To mitigate these issues, the authors propose Symmetric Nucleus Sampling (SNS) to refine training pairs and introduce an Expert Embedding Engine (EEE) that fuses representations from multiple experts. Furthermore, a bias-aware objective function and a projection network are jointly optimized to enhance embedding learning. The proposed approach substantially narrows the modality gap, achieving an average reduction of over 90%. Empirical evaluation demonstrates that the resulting embeddings significantly outperform those generated by hierarchical sampling and conventional curation baselines across downstream tasks.
This study addresses the embedding space incompatibility arising from iterative updates of multimodal models, where full index recomputation incurs prohibitive costs. To mitigate this, the work proposes a Spherical Linear Interpolation (SLERP) strategy leveraging geometric properties to optimize query vectors across legacy and updated models. By integrating contrastive learning with orthogonal posterior alignment, the authors theoretically demonstrate that their approach significantly reduces residual angular error compared to conventional orthogonal alignment. The primary contribution lies in enhancing retrieval compatibility without reconstructing the image gallery. Extensive evaluations across multiple benchmarks reveal substantial improvements in Recall@K, effectively restoring backward compatibility while circumventing the computational overhead associated with large-scale re-indexing.
This work addresses the systematic geometric misalignment—commonly referred to as the modality gap—between visual and linguistic representations in multimodal large language models, which hinders effective alignment and scalability. The authors propose a theoretical framework that decomposes the modality gap within a fixed reference frame, thereby relaxing the isotropic assumption inherent in conventional contrastive learning. This enables, for the first time, statistical alignment using only large-scale unpaired data without requiring aligned image–text pairs. Building on this theory, they introduce ReAlign, a training-free three-step alignment strategy (Anchor-Trace-Centroid), and ReVision, a scalable pretraining paradigm. Experiments demonstrate that the proposed approach significantly improves representation alignment without relying on high-quality paired data, offering a novel pathway toward efficient scaling of multimodal models.
This study addresses the inconsistent impact of the modality gap in vision-language models on downstream tasks by proposing a unified geometric interpretation framework. Building upon CLIP and SigLIP encoders, this work employs rank-one decomposition and geometric analysis of similarity scores to elucidate the geometric nature of the modality gap and its underlying mechanisms across different readout strategies. Our findings reveal that the dominant direction captures between 94.4% and 99.9% of the separation norm, thereby providing a unified explanation for the differential effects of gap modification on tasks such as classification and retrieval. Consequently, this framework offers a solid theoretical foundation for selecting appropriate gap intervention strategies.
本文提出PACE框架,通过两阶段方法逐步扩展表示和可训练参数空间,解决直接优化点积相似度导致的不稳定训练问题。