Score
Designs and trains models that map inputs to a shared embedding space—typically using siamese/twin-network architectures and pairwise similarity losses—so that distances or similarity scores between embeddings reflect labeled pairwise similarity and can be used for matching, alignment, retrieval, and anomaly detection via thresholding. Builds and evaluates the end-to-end similarity pipeline, including similarity scoring and thresholding, efficient/approximate vector search and indexing, embedding evaluation and analysis, and techniques to increase inter-class margins or reduce intra-class variance.
This work addresses the challenge that similarity scores from different embedding models are often incomparable due to geometric discrepancies, which hinders the transferability of fixed similarity thresholds across models. To overcome this limitation without requiring real queries, the authors propose a synthetic query probing method that generates controllable query–text pairs to enable large-scale analysis of cross-model similarity distributions. By learning mappings between score spaces, the approach aligns outputs using calibration strategies including linear regression, isotonic regression, and quantile mapping. Experimental results reveal that while models exhibit consistent ranking behavior, their similarity scores suffer from systematic offsets. The learned mappings substantially improve threshold portability across models, with isotonic regression yielding the best performance.
Embedding models trained independently on similar data capture stable semantic meanings but yield inconsistent representation spaces, hindering interoperability across models. This work addresses compatibility challenges in multimodal search and model upgrades via orthogonal transformation-based embedding alignment. Theoretically, we derive the first tight Procrustes alignment error bound, proving the existence of a near-isometric orthogonal transformation that approximately preserves pairwise inner products—establishing rigorous theoretical foundations for alignment. Methodologically, we employ efficient Procrustes analysis as a post-hoc alignment procedure, preserving the intrinsic geometric structure of each embedding space while enabling cross-model alignment. Experiments demonstrate substantial improvements in model retraining compatibility, text retrieval fusion accuracy, and cross-modal search performance; our method achieves state-of-the-art results in hybrid multimodal search.
This work addresses the weak interpretability of text embedding spaces and their limited structural representation. We propose the Unified Topological Signature (UTS) framework—the first systematic approach to jointly model the topological and geometric structure of embedding spaces. UTS integrates multi-dimensional features, including persistent homology, curvature estimation, and local density, overcoming the redundancy and low discriminability of conventional metrics. By applying clustering analysis and correlation modeling, UTS decodes the mapping between spatial organization and downstream retrieval performance, establishing a quantitative relationship between topological features and document retrievability. Extensive evaluation across multiple state-of-the-art embedding models and benchmark datasets demonstrates that UTS stably predicts inter-model performance differences and ranking effectiveness, exhibiting strong generalization capability and cross-model comparability.
Data selection for large language model (LLM) pretraining relies heavily on embedding-based similarity metrics, yet existing methods lack systematic evaluation grounded in pretraining objectives. Method: We propose the first pretraining-aware evaluation framework for data embedding models, establishing a quantitative linkage between embedding-space similarity and downstream pretraining loss reduction. Our empirical study employs a 1.7B-parameter decoder-only model trained on The Pile. Contribution/Results: We find that simple average word embeddings achieve performance comparable to state-of-the-art complex embedding models—revealing a structural misalignment between current embedding designs and pretraining objectives. The framework not only delineates the efficacy boundaries of diverse embedding methods but also provides, for the first time, a reproducible evaluation paradigm and principled design guidance for developing “pretraining-aware” customized embedding models.
This work proposes a novel approach to large-scale retrieval that circumvents the prohibitive cost of full reranking by constructing query and item embeddings derived from the outputs of a reranker. Specifically, it leverages relevance scores assigned by a heavyweight reranker over a set of support items to generate lightweight embeddings, thereby enabling the reranking model to directly guide embedding learning—a capability demonstrated here for the first time. Under mild conditions, the method is theoretically shown to approximate arbitrarily complex similarity functions. Through systematic investigation of support item selection strategies and integration with approximate nearest neighbor search, the approach significantly improves candidate set quality across multiple academic and industrial datasets while maintaining computational efficiency.
This work investigates whether multi-vector embeddings are strictly more expressive than single-vector embeddings, particularly whether the former can be effectively approximated by the latter under identical representation budgets. By constructing hard instances based on Chamfer similarity—specifically, pattern matrices encoding NANDₖ Boolean functions—and leveraging the Pattern Matrix Method for complexity analysis, the paper establishes the first rigorous proof that no low-dimensional single-vector embedding can linearly approximate the similarity captured by multi-vector embeddings. Moreover, it demonstrates that any single-vector embedding approximating such a multi-vector representation requires dimensionality at least (ε²m)^Ω(1/ε), thereby revealing a fundamental gap in representational efficiency between the two paradigms.
This work addresses the challenge of setting an appropriate distance threshold in Siamese verification networks by proposing an unsupervised method for automatic threshold determination. Relying on the assumption that the distribution of embedding distances exhibits a bimodal structure, the method dynamically identifies the local minimum between the two modes to establish the verification threshold. It requires no labeled data and supports real-time updates in deployment environments. Experimental results demonstrate that the proposed approach achieves an average verification accuracy of 94% across four benchmark datasets—MNIST, CIFAR-10, LFW, and PKLot—matching the performance of supervised equal error rate (EER)-based methods while substantially reducing the need for costly manual annotation.