siamese metric learning

Designs and trains models that map inputs to a shared embedding space—typically using siamese/twin-network architectures and pairwise similarity losses—so that distances or similarity scores between embeddings reflect labeled pairwise similarity and can be used for matching, alignment, retrieval, and anomaly detection via thresholding. Builds and evaluates the end-to-end similarity pipeline, including similarity scoring and thresholding, efficient/approximate vector search and indexing, embedding evaluation and analysis, and techniques to increase inter-class margins or reduce intra-class variance.

siamesemetriclearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$182K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge that similarity scores from different embedding models are often incomparable due to geometric discrepancies, which hinders the transferability of fixed similarity thresholds across models. To overcome this limitation without requiring real queries, the authors propose a synthetic query probing method that generates controllable query–text pairs to enable large-scale analysis of cross-model similarity distributions. By learning mappings between score spaces, the approach aligns outputs using calibration strategies including linear regression, isotonic regression, and quantile mapping. Experimental results reveal that while models exhibit consistent ranking behavior, their similarity scores suffer from systematic offsets. The learned mappings substantially improve threshold portability across models, with isotonic regression yielding the best performance.

cross-model comparabilityembedding modelsretrieval-augmented generation

When Embedding Models Meet: Procrustes Bounds and Applications

Oct 15, 2025
LM
Lucas Maystre
🏛️ UiPath | Spotify

Embedding models trained independently on similar data capture stable semantic meanings but yield inconsistent representation spaces, hindering interoperability across models. This work addresses compatibility challenges in multimodal search and model upgrades via orthogonal transformation-based embedding alignment. Theoretically, we derive the first tight Procrustes alignment error bound, proving the existence of a near-isometric orthogonal transformation that approximately preserves pairwise inner products—establishing rigorous theoretical foundations for alignment. Methodologically, we employ efficient Procrustes analysis as a post-hoc alignment procedure, preserving the intrinsic geometric structure of each embedding space while enabling cross-model alignment. Experiments demonstrate substantial improvements in model retraining compatibility, text retrieval fusion accuracy, and cross-modal search performance; our method achieves state-of-the-art results in hybrid multimodal search.

Aligning embeddings from separately trained modelsEnabling interoperability through orthogonal transformationsImproving multimodal search and model compatibility

From Topology to Retrieval: Decoding Embedding Spaces with Unified Signatures

Nov 27, 2025
FR
Florian Rottach
🏛️ University of Tübingen | The University of Texas at Austin | Fribourg University

This work addresses the weak interpretability of text embedding spaces and their limited structural representation. We propose the Unified Topological Signature (UTS) framework—the first systematic approach to jointly model the topological and geometric structure of embedding spaces. UTS integrates multi-dimensional features, including persistent homology, curvature estimation, and local density, overcoming the redundancy and low discriminability of conventional metrics. By applying clustering analysis and correlation modeling, UTS decodes the mapping between spatial organization and downstream retrieval performance, establishing a quantitative relationship between topological features and document retrievability. Extensive evaluation across multiple state-of-the-art embedding models and benchmark datasets demonstrates that UTS stably predicts inter-model performance differences and ranking effectiveness, exhibiting strong generalization capability and cross-model comparability.

Analyzing topological and geometric measures of text embedding spacesIntroducing a unified framework to characterize embedding spaces holisticallyLinking topological structure to retrieval performance and model properties

Analyzing Similarity Metrics for Data Selection for Language Model Pretraining

Feb 04, 2025
DS
Dylan Sam
🏛️ Carnegie Mellon University | Google

Data selection for large language model (LLM) pretraining relies heavily on embedding-based similarity metrics, yet existing methods lack systematic evaluation grounded in pretraining objectives. Method: We propose the first pretraining-aware evaluation framework for data embedding models, establishing a quantitative linkage between embedding-space similarity and downstream pretraining loss reduction. Our empirical study employs a 1.7B-parameter decoder-only model trained on The Pile. Contribution/Results: We find that simple average word embeddings achieve performance comparable to state-of-the-art complex embedding models—revealing a structural misalignment between current embedding designs and pretraining objectives. The framework not only delineates the efficacy boundaries of diverse embedding methods but also provides, for the first time, a reproducible evaluation paradigm and principled design guidance for developing “pretraining-aware” customized embedding models.

Analyzing embedding models for data curationEvaluating embedding models for language modelsMeasuring similarity impact on pretraining loss

Latest Papers

What's happening recently
View more

This work proposes a novel approach to large-scale retrieval that circumvents the prohibitive cost of full reranking by constructing query and item embeddings derived from the outputs of a reranker. Specifically, it leverages relevance scores assigned by a heavyweight reranker over a set of support items to generate lightweight embeddings, thereby enabling the reranking model to directly guide embedding learning—a capability demonstrated here for the first time. Under mild conditions, the method is theoretically shown to approximate arbitrarily complex similarity functions. Through systematic investigation of support item selection strategies and integration with approximate nearest neighbor search, the approach significantly improves candidate set quality across multiple academic and industrial datasets while maintaining computational efficiency.

candidate retrievalembeddingranking

This work investigates whether multi-vector embeddings are strictly more expressive than single-vector embeddings, particularly whether the former can be effectively approximated by the latter under identical representation budgets. By constructing hard instances based on Chamfer similarity—specifically, pattern matrices encoding NANDₖ Boolean functions—and leveraging the Pattern Matrix Method for complexity analysis, the paper establishes the first rigorous proof that no low-dimensional single-vector embedding can linearly approximate the similarity captured by multi-vector embeddings. Moreover, it demonstrates that any single-vector embedding approximating such a multi-vector representation requires dimensionality at least (ε²m)^Ω(1/ε), thereby revealing a fundamental gap in representational efficiency between the two paradigms.

approximation hardnessChamfer similaritymulti-vector embeddings

This work addresses the challenge of setting an appropriate distance threshold in Siamese verification networks by proposing an unsupervised method for automatic threshold determination. Relying on the assumption that the distribution of embedding distances exhibits a bimodal structure, the method dynamically identifies the local minimum between the two modes to establish the verification threshold. It requires no labeled data and supports real-time updates in deployment environments. Experimental results demonstrate that the proposed approach achieves an average verification accuracy of 94% across four benchmark datasets—MNIST, CIFAR-10, LFW, and PKLot—matching the performance of supervised equal error rate (EER)-based methods while substantially reducing the need for costly manual annotation.

distance thresholdembedding spaceSiamese networks

Hot Scholars

TL

Trung Le

Faculty of Information Technology, Monash University, Australia
Adversarial Machine LearningGenerative ModelsModel UnlearningModel Editing
XZ

Xiangyu Zhao

Associate Professor, City University of Hong Kong
RecommendationsLarge Language Models (LLMs)TrustworthyAISearch Engine
BB

Binod Bhattarai

Assistant Professor, University of Aberdeen
Machine LearningMedical Image AnalysisComputer Vision
JY

Junchi Yan

FIAPR & ICML Board Member, SJTU (2018-), SII (2024-), AWS (2019-2022), IBM (2011-2018)
Computational IntelligenceAI4ScienceMachine LearningAutonomous Driving
CH

Conghui He

Shanghai AI Laboratory
Data-centric AILLMDocument Intelligence