π€ AI Summary
This study addresses the embedding space incompatibility arising from iterative updates of multimodal models, where full index recomputation incurs prohibitive costs. To mitigate this, the work proposes a Spherical Linear Interpolation (SLERP) strategy leveraging geometric properties to optimize query vectors across legacy and updated models. By integrating contrastive learning with orthogonal posterior alignment, the authors theoretically demonstrate that their approach significantly reduces residual angular error compared to conventional orthogonal alignment. The primary contribution lies in enhancing retrieval compatibility without reconstructing the image gallery. Extensive evaluations across multiple benchmarks reveal substantial improvements in Recall@K, effectively restoring backward compatibility while circumventing the computational overhead associated with large-scale re-indexing.
π Abstract
Contrastive vision-language models map visual and textual representations into a shared normalized embedding space, making cosine similarity the natural metric for cross-modal retrieval. A practical challenge arises during model upgrades: independently trained models generally produce incompatible representation spaces, so replacing a deployed model typically requires recomputing embeddings for the entire gallery, which is prohibitively expensive at scale. Orthogonal post-hoc alignment can partially mitigate this problem by mapping new-model queries into the old-model gallery space. However, because independently trained models can differ in fine-grained representation structure, the orthogonal alignment remains approximate, leaving a residual angular discrepancy between the old-model query and the aligned new-model query. We study whether interpolation along the spherical geodesic between these two normalized query representations can improve retrieval without re-indexing the gallery. We characterize when this path contains an interior query direction closer to an idealized retrieval-optimal direction than either endpoint, and connect this characterization to Recall@$K$ through a local margin-based certification result. Experiments across multiple benchmarks and model families show that post-alignment spherical interpolation improves over orthogonal alignment alone, recovering backward-compatibility in most evaluated settings. Consistent with our geometric characterization, per-query oracle analysis shows that retrieval-favorable interior points occur frequently in practice. Code is available at https://github.com/miccunifi/SLERP_backward_compatibility .