embedding retrieval

Designs and builds vector-based retrieval systems that index, store, and search embedding representations (e.g., text embeddings) to return semantically similar chunks or documents. Implements nearest-neighbor/ANN indexes, similarity scoring and reranking pipelines, and selection or filtering of top results for downstream consumers such as LLM context windows or other components.

embeddingretrieval

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$184K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a semantic similarity computation method that integrates Word Mover’s Distance (WMD) with pretrained word embeddings such as GloVe to better model the semantic relationship between queries and documents in information retrieval. Traditional centroid-based word embedding approaches often fail to capture fine-grained semantic matches, particularly when handling synonymy and polysemy. By minimizing the transportation cost of aligning query and document terms in the embedding space, the proposed method achieves a more precise representation of semantic correspondence. Experimental results demonstrate that this approach significantly outperforms baseline models—including Doc2Vec and Latent Semantic Analysis (LSA)—on similarity ranking tasks, while maintaining domain independence and high retrieval accuracy, thereby confirming its effectiveness and generalizability in practical information retrieval scenarios.

distributional semanticsinformation retrievalquery similarity

To address the significant degradation in retrieval accuracy of vector similarity search under complex semantic queries—such as those involving constraints, negation, or abstract concepts—this paper proposes a two-stage retrieval framework: an efficient initial retrieval using approximate nearest neighbor (ANN) algorithms (e.g., FAISS), followed by context-aware fine-grained re-ranking powered by large language models (LLMs). Distinct from prior approaches, our work is the first to deeply integrate LLMs into the vector search pipeline, leveraging customized prompt engineering and a structured evaluation framework to achieve precise semantic understanding of complex queries while maintaining millisecond-scale latency. Experimental results across multiple structured benchmarks demonstrate that our method improves accuracy by 32% on average over baseline vector-only search.

Complex Query UnderstandingInformation RetrievalVector Similarity Search

Beyond Nearest Neighbors: Semantic Compression and Graph-Augmented Retrieval for Enhanced Vector Search

Jul 25, 2025
RR
Rahul Raja
🏛️ Carnegie Mellon University | LinkedIn | Stanford University | Boston University

Traditional vector retrieval relies on approximate nearest neighbor (ANN) search, which often yields semantically redundant results and fails to meet the diversity and contextual richness requirements of applications such as retrieval-augmented generation (RAG) and multi-hop question answering. To address this, we propose a novel paradigm—“Semantic Compression and Graph-Enhanced Retrieval”—that introduces submodular optimization into vector retrieval for the first time. We formalize semantic compression to maximize information coverage while explicitly suppressing redundancy. Leveraging information-geometric similarity metrics and k-nearest neighbor (kNN) graphs, we construct a multi-hop semantic search framework, further augmented with knowledge graph integration for structured semantic querying. Our method supports hybrid indexing and significantly improves semantic diversity and coverage in high-dimensional embedding spaces, outperforming state-of-the-art ANN baselines. The implementation is open-sourced, advancing research toward semantics-centric vector search.

Addresses redundancy in nearest neighbor search via coverage optimizationEnhances vector search with semantic compression for diversityIntroduces graph-augmented retrieval for context-aware multi-hop search

This study addresses the lack of systematic evaluation of hybrid search mechanisms that combine semantic retrieval with metadata filtering in existing vector databases. We propose a novel relevance metric, Global-Local Selectivity (GLS), construct MoReVec—the first benchmark dataset supporting filtered retrieval—and extend ANN-Benchmarks to enable unified evaluation of hybrid search performance. Through comprehensive experiments integrating diverse filtering strategies into FAISS, Milvus, and pgvector with IVFFlat and HNSW indexes, we demonstrate that engine-level algorithmic integration critically governs performance: Milvus achieves more stable recall via hybrid execution, pgvector’s optimizer often selects suboptimal query plans, and IVFFlat outperforms HNSW under low-selectivity queries. Our findings culminate in practical configuration guidelines that offer both theoretical insights and actionable recommendations for efficient hybrid search deployment.

Filtered Approximate Nearest Neighbor SearchFiltering StrategiesHybrid Search

pEBR: A Probabilistic Approach to Embedding Based Retrieval

Oct 25, 2024
HZ
Han Zhang
🏛️ JD.com | Meta | Shanghai Jiaotong University

Existing embedding-based retrieval systems employ fixed recall budgets, leading to insufficient recall for head queries and degraded precision for tail queries—rooted in the frequentist paradigm of loss functions. This paper proposes a query-aware probabilistic retrieval framework that abandons fixed truncation and instead models the query-specific distribution of candidate cosine similarities. By learning a query-conditioned cumulative distribution function (CDF), the framework dynamically determines similarity thresholds per query. For the first time, it jointly optimizes precision and recall for both head and tail queries within a unified probabilistic framework. Experiments across multiple industrial retrieval benchmarks demonstrate significant improvements in overall effectiveness. Ablation studies validate the efficacy of explicitly modeling head–tail disparities. The core innovation lies in elevating threshold selection from heuristic, static configuration to a principled, query-adaptive, probabilistic decision process.

Addresses fixed retrieval cutoff limitations in embedding systemsImproves recall for head queries and precision for tail queriesLearns probability distributions of relevant items for adaptive retrieval

Latest Papers

What's happening recently
View more

This work proposes a novel approach to large-scale retrieval that circumvents the prohibitive cost of full reranking by constructing query and item embeddings derived from the outputs of a reranker. Specifically, it leverages relevance scores assigned by a heavyweight reranker over a set of support items to generate lightweight embeddings, thereby enabling the reranking model to directly guide embedding learning—a capability demonstrated here for the first time. Under mild conditions, the method is theoretically shown to approximate arbitrarily complex similarity functions. Through systematic investigation of support item selection strategies and integration with approximate nearest neighbor search, the approach significantly improves candidate set quality across multiple academic and industrial datasets while maintaining computational efficiency.

candidate retrievalembeddingranking

This work addresses the gap between benchmark-driven embedding model selection and real-world deployment constraints by introducing the first framework to evaluate embedding models within a complete retrieval pipeline. It systematically compares the end-to-end performance of T3EM’s commercial API against leading open-source models across diverse tasks—including retrieval, classification, clustering, and semantic similarity—as covered by the MTEB benchmark, while jointly accounting for latency, cost, task type, and deployment conditions. The study develops a comprehensive, end-to-end model selection guide encompassing embedding generation, indexing, search, and chunking strategies, revealing significant performance discrepancies that emerge only in full-system contexts. These insights provide practitioners with actionable, empirically grounded criteria for embedding model adoption in real-world applications.

deployment constraintsmodel selectionpractical benchmarking

This study addresses the lack of comprehensive, reproducible benchmarks for vector databases across retrieval quality, latency, throughput, and resource consumption. For the first time, it jointly evaluates seven systems—FAISS, Qdrant, Milvus, Weaviate, Chroma, pgvector, and LanceDB—across six datasets (including SIFT, GIST, MS MARCO, and GloVe) comprising over four million vectors, measuring fifteen performance and resource metrics. The evaluation reveals that FAISS achieves the highest single-node throughput (866 QPS), Weaviate delivers out-of-the-box recall exceeding 99%, Qdrant exhibits the lowest median latency (4.55 ms), and LanceDB constructs indexes fastest albeit with slightly lower retrieval quality. The authors open-source the complete benchmarking framework and propose practical guidelines for system selection.

approximate nearest neighbor searchquery latencyresource utilization

Hot Scholars

DE

Diego Elias Costa

Assistant Professor, Concordia University
Software EngineeringSoftware EcosystemsPerformance EngineeringSE4AI
YT

Yicheng Tao

Carnegie Mellon Univeristy
Natural Language ProcessingSmart Cities
RY

Ruifeng Yuan

Ph.D from the Hong Kong Polytechnic University
Nature language processing
TX

Tianyi Xu

Tulane University
Reinforcement LearningNetwork OptimizaitonStatisticsNLP(LLM)
XZ

Xiaoyong Zhu

Jiangsu University
Electrical MachinesElectrical Vehicle