Score
Designs and implements methods and systems to compute semantic similarity between pieces of text or their vector/embedding representations, including metrics, embedding-based scoring, retrieval and search algorithms, and procedures for aligning semantically related segments. Builds evaluation protocols and analyses that compare similarity measures against human judgments, tune metrics, and assess retrieval/search performance.
This work addresses the joint optimization of large language models (LLMs) and text embedding techniques to enhance efficiency and robustness in semantic matching, clustering, and information retrieval. We propose the first unified taxonomy centered on the *interaction patterns* between LLMs and embeddings—categorizing approaches into three paradigms: LLM-augmented embeddings, LLM-as-embedder, and LLM-understanding-embeddings—thereby transcending conventional task-centric taxonomies. By integrating supervised/unsupervised embedding learning, instruction tuning, prompt engineering, representation space analysis, and interpretability methods, we construct a structured knowledge graph encompassing over 100 studies. Our framework precisely delineates capability boundaries and application scopes for each paradigm, identifies persistent limitations inherited from pre-trained language models (PLMs) and novel challenges introduced by LLMs, and provides a theoretically grounded, empirically informed roadmap for future advancement.
This work proposes a semantic similarity computation method that integrates Word Mover’s Distance (WMD) with pretrained word embeddings such as GloVe to better model the semantic relationship between queries and documents in information retrieval. Traditional centroid-based word embedding approaches often fail to capture fine-grained semantic matches, particularly when handling synonymy and polysemy. By minimizing the transportation cost of aligning query and document terms in the embedding space, the proposed method achieves a more precise representation of semantic correspondence. Experimental results demonstrate that this approach significantly outperforms baseline models—including Doc2Vec and Latent Semantic Analysis (LSA)—on similarity ranking tasks, while maintaining domain independence and high retrieval accuracy, thereby confirming its effectiveness and generalizability in practical information retrieval scenarios.
This study addresses the fundamental question of whether semantic similarity measures genuinely comprehend semantic relationships. We propose the first evaluation framework based on controlled, small-scale semantic transformations to systematically assess the semantic discrimination capability of 18 state-of-the-art methods—including bag-of-words, embedding-based, LLM-based, and structure-aware models—on software engineering texts and code. Experiments reveal that mainstream embedding methods exhibit up to 99.9% misclassification rates in semantic opposition scenarios, exposing their reliance on superficial surface patterns. Substituting cosine similarity for Euclidean distance improves performance by 24–66%. LLM-based methods demonstrate superior fine-grained semantic distinction. Critically, our framework uncovers a foundational limitation in existing measures: their failure to capture semantic essence. It establishes the first reproducible, scalable benchmark paradigm for trustworthy semantic computation in software engineering contexts.
This study addresses the cross-modal retrieval challenge for multimodal educational content—particularly computer science textbooks containing interleaved text and figures. We propose a multi-vector representation method that jointly encodes textual and visual semantics using a vision-language model (VLM), generating fine-grained multimodal embeddings and indexing them in a vector database to enable efficient cross-modal retrieval. Evaluated on over 3,600 pages of textbook material, we systematically compare four similarity metrics and find cosine similarity significantly outperforms alternatives. Benchmarking against 75 natural-language queries confirms substantial improvements in retrieval precision and practical utility within digital library settings. Our approach delivers a reproducible, scalable technical framework for intelligent discovery of multimodal educational resources, advancing the state of cross-modal semantic search in academic and pedagogical contexts.
To address the significant degradation in retrieval accuracy of vector similarity search under complex semantic queries—such as those involving constraints, negation, or abstract concepts—this paper proposes a two-stage retrieval framework: an efficient initial retrieval using approximate nearest neighbor (ANN) algorithms (e.g., FAISS), followed by context-aware fine-grained re-ranking powered by large language models (LLMs). Distinct from prior approaches, our work is the first to deeply integrate LLMs into the vector search pipeline, leveraging customized prompt engineering and a structured evaluation framework to achieve precise semantic understanding of complex queries while maintaining millisecond-scale latency. Experimental results across multiple structured benchmarks demonstrate that our method improves accuracy by 32% on average over baseline vector-only search.
To address insufficient relevance in Top-N results for semantic document retrieval, this paper proposes a semantic matching framework integrated with evolutionary optimization. First, dense semantic vectors for queries and documents are generated using the Universal Sentence Encoder. Then, genetic algorithms (GA) and differential evolution (DE) are jointly and innovatively incorporated into the similarity ranking process to end-to-end optimize the matching function within the semantic space. This work is the first to systematically integrate both GA and DE into a semantic retrieval framework, thereby significantly enhancing ranking robustness and accuracy. Experimental evaluation on the SQuAD dataset demonstrates substantial improvements in Top-N accuracy over conventional static distance metrics—such as Manhattan distance—validating the effectiveness of evolutionary optimization in strengthening semantic matching performance.
This work addresses a critical limitation in traditional approximate nearest neighbor (ANN) retrieval evaluation, which relies solely on recall and fails to distinguish semantically relevant from irrelevant neighbors, often misrepresenting retrieval quality. To overcome this, the authors introduce Semantic Recall—a novel metric that quantifies only those semantically relevant results theoretically recoverable via exact search—and propose Tolerant Recall as an efficient proxy for practical evaluation. This is the first systematic integration of semantic relevance into vector retrieval assessment, revealing the pervasive sparsity of relevant results within embedding spaces. Experimental results demonstrate that the proposed metrics more accurately reflect real-world retrieval effectiveness, and algorithms optimized under this framework achieve superior trade-offs between cost and quality.
This work addresses the gap between benchmark-driven embedding model selection and real-world deployment constraints by introducing the first framework to evaluate embedding models within a complete retrieval pipeline. It systematically compares the end-to-end performance of T3EM’s commercial API against leading open-source models across diverse tasks—including retrieval, classification, clustering, and semantic similarity—as covered by the MTEB benchmark, while jointly accounting for latency, cost, task type, and deployment conditions. The study develops a comprehensive, end-to-end model selection guide encompassing embedding generation, indexing, search, and chunking strategies, revealing significant performance discrepancies that emerge only in full-system contexts. These insights provide practitioners with actionable, empirically grounded criteria for embedding model adoption in real-world applications.
This study systematically investigates the intrinsic semantic and syntactic properties of mainstream word embedding methods—such as Word2Vec and GloVe—and their performance disparities across diverse natural language processing tasks. By establishing a unified evaluation framework that integrates publicly available pretrained embeddings with standard benchmark datasets, the work conducts empirical comparisons on canonical tasks including semantic similarity and analogical reasoning. The findings delineate the performance boundaries and optimal application scenarios for each embedding model, offering practitioners reliable guidance for model selection in real-world settings. Furthermore, the analysis deepens the understanding of the inherent limitations of static word representations, highlighting critical constraints in capturing contextual and compositional linguistic phenomena.
This work addresses the challenge that similarity scores from different embedding models are often incomparable due to geometric discrepancies, which hinders the transferability of fixed similarity thresholds across models. To overcome this limitation without requiring real queries, the authors propose a synthetic query probing method that generates controllable query–text pairs to enable large-scale analysis of cross-model similarity distributions. By learning mappings between score spaces, the approach aligns outputs using calibration strategies including linear regression, isotonic regression, and quantile mapping. Experimental results reveal that while models exhibit consistent ranking behavior, their similarity scores suffer from systematic offsets. The learned mappings substantially improve threshold portability across models, with isotonic regression yielding the best performance.
This study addresses the problem of semantic entanglement in vector retrieval, where multi-topic documents induce overlapping semantics in embedding space, degrading retrieval accuracy. The work formally defines semantic entanglement for the first time and introduces a quantifiable Entanglement Index (EI). To mitigate this issue, the authors propose a context-conditioned, four-stage Semantic Disentanglement Pipeline (SDP) that dynamically optimizes pre-embedding text organization through document restructuring and an agent-based feedback mechanism. Experimental evaluation on over 2,000 medical documents demonstrates that the approach substantially alleviates semantic entanglement, reducing the average EI from 0.71 to 0.14 and improving Top-K retrieval accuracy from 32% to 82%.