Mapping Similarity Spaces across Embedding Models with Synthetic Query Probing

📅 2026-08-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that similarity scores from different embedding models are often incomparable due to geometric discrepancies, which hinders the transferability of fixed similarity thresholds across models. To overcome this limitation without requiring real queries, the authors propose a synthetic query probing method that generates controllable query–text pairs to enable large-scale analysis of cross-model similarity distributions. By learning mappings between score spaces, the approach aligns outputs using calibration strategies including linear regression, isotonic regression, and quantile mapping. Experimental results reveal that while models exhibit consistent ranking behavior, their similarity scores suffer from systematic offsets. The learned mappings substantially improve threshold portability across models, with isotonic regression yielding the best performance.
📝 Abstract
Retrieval-Augmented Generation systems rely on similarity scores to retrieve relevant content, yet scores are not directly comparable across embedding models due to differing geometric properties, complicating model migration and limiting threshold reuse. We study how similarity scores can be related by learning mappings between score distributions rather than embeddings. We introduce Synthetic Query Probing, generating queries from documents to create controlled query-chunk pairs, enabling large-scale, reference-free analysis of cross-model similarity behavior. We evaluate the approach on multiple embedding configurations and learn score conversion functions using linear, isotonic, and quantile mappings. Experiments on SciFact and a proprietary corpus show that while models largely agree on rankings, their absolute scores exhibit systematic distortions. Learned mappings partially align these spaces and improve threshold portability, with isotonic regression performing best. Our results highlight the need for cross-model calibration and position Synthetic Query Probing as a scalable framework for analyzing embedding comparability.
Problem

Research questions and friction points this paper is trying to address.

embedding models
similarity scores
cross-model comparability
threshold portability
retrieval-augmented generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synthetic Query Probing
Similarity Score Calibration
Cross-Model Embedding Alignment
Isotonic Regression
Retrieval-Augmented Generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Marcin Rozmus
GenAI Engineering, Pegasystems, Kraków, Poland
P
Peter van der Putten
AI Lab, Pegasystems, Amsterdam, The Netherlands; LIACS, Leiden University, Leiden, the Netherlands