measure semantic similarity

Designs and implements methods and systems to compute semantic similarity between pieces of text or their vector/embedding representations, including metrics, embedding-based scoring, retrieval and search algorithms, and procedures for aligning semantically related segments. Builds evaluation protocols and analyses that compare similarity measures against human judgments, tune metrics, and assess retrieval/search performance.

measuresemanticsimilarity

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a semantic similarity computation method that integrates Word Mover’s Distance (WMD) with pretrained word embeddings such as GloVe to better model the semantic relationship between queries and documents in information retrieval. Traditional centroid-based word embedding approaches often fail to capture fine-grained semantic matches, particularly when handling synonymy and polysemy. By minimizing the transportation cost of aligning query and document terms in the embedding space, the proposed method achieves a more precise representation of semantic correspondence. Experimental results demonstrate that this approach significantly outperforms baseline models—including Doc2Vec and Latent Semantic Analysis (LSA)—on similarity ranking tasks, while maintaining domain independence and high retrieval accuracy, thereby confirming its effectiveness and generalizability in practical information retrieval scenarios.

distributional semanticsinformation retrievalquery similarity

How Small Transformation Expose the Weakness of Semantic Similarity Measures

Sep 08, 2025
SL
Serge Lionel Nikiema
🏛️ University of Luxembourg

This study addresses the fundamental question of whether semantic similarity measures genuinely comprehend semantic relationships. We propose the first evaluation framework based on controlled, small-scale semantic transformations to systematically assess the semantic discrimination capability of 18 state-of-the-art methods—including bag-of-words, embedding-based, LLM-based, and structure-aware models—on software engineering texts and code. Experiments reveal that mainstream embedding methods exhibit up to 99.9% misclassification rates in semantic opposition scenarios, exposing their reliance on superficial surface patterns. Substituting cosine similarity for Euclidean distance improves performance by 24–66%. LLM-based methods demonstrate superior fine-grained semantic distinction. Critically, our framework uncovers a foundational limitation in existing measures: their failure to capture semantic essence. It establishes the first reproducible, scalable benchmark paradigm for trustworthy semantic computation in software engineering contexts.

Evaluating semantic similarity measures for software engineering tasksIdentifying flaws where methods confuse opposites and synonymsTesting 18 methods including embeddings and LLMs on semantic understanding

Vector embedding of multi-modal texts: a tool for discovery?

Sep 09, 2025
BP
Beth Plale
🏛️ University of Oregon | Indiana University

This study addresses the cross-modal retrieval challenge for multimodal educational content—particularly computer science textbooks containing interleaved text and figures. We propose a multi-vector representation method that jointly encodes textual and visual semantics using a vision-language model (VLM), generating fine-grained multimodal embeddings and indexing them in a vector database to enable efficient cross-modal retrieval. Evaluated on over 3,600 pages of textbook material, we systematically compare four similarity metrics and find cosine similarity significantly outperforms alternatives. Benchmarking against 75 natural-language queries confirms substantial improvements in retrieval precision and practical utility within digital library settings. Our approach delivers a reproducible, scalable technical framework for intelligent discovery of multimodal educational resources, advancing the state of cross-modal semantic search in academic and pedagogical contexts.

Benchmarking retrieval performance across different similarity measuresExploring strengths and weaknesses of vision-language models for retrievalImproving discovery in multi-modal content using vector embeddings

To address the significant degradation in retrieval accuracy of vector similarity search under complex semantic queries—such as those involving constraints, negation, or abstract concepts—this paper proposes a two-stage retrieval framework: an efficient initial retrieval using approximate nearest neighbor (ANN) algorithms (e.g., FAISS), followed by context-aware fine-grained re-ranking powered by large language models (LLMs). Distinct from prior approaches, our work is the first to deeply integrate LLMs into the vector search pipeline, leveraging customized prompt engineering and a structured evaluation framework to achieve precise semantic understanding of complex queries while maintaining millisecond-scale latency. Experimental results across multiple structured benchmarks demonstrate that our method improves accuracy by 32% on average over baseline vector-only search.

Complex Query UnderstandingInformation RetrievalVector Similarity Search

Evolutionary Algorithms Approach For Search Based On Semantic Document Similarity

Aug 04, 2023
CM
Chandrashekar Muniyappa
🏛️ University of North Dakota

To address insufficient relevance in Top-N results for semantic document retrieval, this paper proposes a semantic matching framework integrated with evolutionary optimization. First, dense semantic vectors for queries and documents are generated using the Universal Sentence Encoder. Then, genetic algorithms (GA) and differential evolution (DE) are jointly and innovatively incorporated into the similarity ranking process to end-to-end optimize the matching function within the semantic space. This work is the first to systematically integrate both GA and DE into a semantic retrieval framework, thereby significantly enhancing ranking robustness and accuracy. Experimental evaluation on the SQuAD dataset demonstrates substantial improvements in Top-N accuracy over conventional static distance metrics—such as Manhattan distance—validating the effectiveness of evolutionary optimization in strengthening semantic matching performance.

Applying evolutionary algorithms for search efficiencyComparing GA and DE with traditional ranking methodsEnhancing document retrieval using semantic similarity

Latest Papers

What's happening recently
View more

This work addresses a critical limitation in traditional approximate nearest neighbor (ANN) retrieval evaluation, which relies solely on recall and fails to distinguish semantically relevant from irrelevant neighbors, often misrepresenting retrieval quality. To overcome this, the authors introduce Semantic Recall—a novel metric that quantifies only those semantically relevant results theoretically recoverable via exact search—and propose Tolerant Recall as an efficient proxy for practical evaluation. This is the first systematic integration of semantic relevance into vector retrieval assessment, revealing the pervasive sparsity of relevant results within embedding spaces. Experimental results demonstrate that the proposed metrics more accurately reflect real-world retrieval effectiveness, and algorithms optimized under this framework achieve superior trade-offs between cost and quality.

Approximate Nearest Neighbor SearchEmbedding DatasetsEvaluation Metrics

This work addresses the gap between benchmark-driven embedding model selection and real-world deployment constraints by introducing the first framework to evaluate embedding models within a complete retrieval pipeline. It systematically compares the end-to-end performance of T3EM’s commercial API against leading open-source models across diverse tasks—including retrieval, classification, clustering, and semantic similarity—as covered by the MTEB benchmark, while jointly accounting for latency, cost, task type, and deployment conditions. The study develops a comprehensive, end-to-end model selection guide encompassing embedding generation, indexing, search, and chunking strategies, revealing significant performance discrepancies that emerge only in full-system contexts. These insights provide practitioners with actionable, empirically grounded criteria for embedding model adoption in real-world applications.

deployment constraintsmodel selectionpractical benchmarking

This study systematically investigates the intrinsic semantic and syntactic properties of mainstream word embedding methods—such as Word2Vec and GloVe—and their performance disparities across diverse natural language processing tasks. By establishing a unified evaluation framework that integrates publicly available pretrained embeddings with standard benchmark datasets, the work conducts empirical comparisons on canonical tasks including semantic similarity and analogical reasoning. The findings delineate the performance boundaries and optimal application scenarios for each embedding model, offering practitioners reliable guidance for model selection in real-world settings. Furthermore, the analysis deepens the understanding of the inherent limitations of static word representations, highlighting critical constraints in capturing contextual and compositional linguistic phenomena.

empirical investigationnatural language processingvector representations

This work addresses the challenge that similarity scores from different embedding models are often incomparable due to geometric discrepancies, which hinders the transferability of fixed similarity thresholds across models. To overcome this limitation without requiring real queries, the authors propose a synthetic query probing method that generates controllable query–text pairs to enable large-scale analysis of cross-model similarity distributions. By learning mappings between score spaces, the approach aligns outputs using calibration strategies including linear regression, isotonic regression, and quantile mapping. Experimental results reveal that while models exhibit consistent ranking behavior, their similarity scores suffer from systematic offsets. The learned mappings substantially improve threshold portability across models, with isotonic regression yielding the best performance.

cross-model comparabilityembedding modelsretrieval-augmented generation

This study addresses the problem of semantic entanglement in vector retrieval, where multi-topic documents induce overlapping semantics in embedding space, degrading retrieval accuracy. The work formally defines semantic entanglement for the first time and introduces a quantifiable Entanglement Index (EI). To mitigate this issue, the authors propose a context-conditioned, four-stage Semantic Disentanglement Pipeline (SDP) that dynamically optimizes pre-embedding text organization through document restructuring and an agent-based feedback mechanism. Experimental evaluation on over 2,000 medical documents demonstrates that the approach substantially alleviates semantic entanglement, reducing the average EI from 0.71 to 0.14 and improving Top-K retrieval accuracy from 32% to 82%.

embedding spaceretrieval precisionRetrieval-Augmented Generation

Hot Scholars

SB

Serge Belongie

University of Copenhagen
Computer VisionMachine Learning
RG

Ross Greer

University of California Merced
Artificial IntelligenceMachine VisionAutonomous DrivingHuman-Robot Interaction
ZJ

Zhi Jin

Sun Yat-Sen University, Associate Professor
MT

Mingkui Tan

South China University of Technology
Machine LearningLarge-scale Optimization
YC

Yiyi Chen

PhD Candidate, Aalborg University
Machine LearningDeep LearningNatural Language Processing