Score
Designs and builds vector-based retrieval systems that index, store, and search embedding representations (e.g., text embeddings) to return semantically similar chunks or documents. Implements nearest-neighbor/ANN indexes, similarity scoring and reranking pipelines, and selection or filtering of top results for downstream consumers such as LLM context windows or other components.
This work addresses the lack of systematic design principles for neural retrieval systems that balance efficiency and effectiveness. It proposes the first vertically layered four-tier framework—spanning representation, granularity, orchestration, and robustness—to structurally characterize key design decisions at each layer and their interdependencies. By integrating Bi- and Cross-encoder architectures, atomic and hierarchical chunking strategies, multi-stage re-ranking, agent-based decomposition, and domain generalization techniques, the study elucidates the mechanistic impact of each design choice on system performance. This approach effectively mitigates critical challenges such as information bottlenecks, semantic blind spots, and temporal drift, thereby offering a practical and actionable optimization pathway for building efficient and robust embedded retrieval systems.
This work proposes a semantic similarity computation method that integrates Word Mover’s Distance (WMD) with pretrained word embeddings such as GloVe to better model the semantic relationship between queries and documents in information retrieval. Traditional centroid-based word embedding approaches often fail to capture fine-grained semantic matches, particularly when handling synonymy and polysemy. By minimizing the transportation cost of aligning query and document terms in the embedding space, the proposed method achieves a more precise representation of semantic correspondence. Experimental results demonstrate that this approach significantly outperforms baseline models—including Doc2Vec and Latent Semantic Analysis (LSA)—on similarity ranking tasks, while maintaining domain independence and high retrieval accuracy, thereby confirming its effectiveness and generalizability in practical information retrieval scenarios.
To address the significant degradation in retrieval accuracy of vector similarity search under complex semantic queries—such as those involving constraints, negation, or abstract concepts—this paper proposes a two-stage retrieval framework: an efficient initial retrieval using approximate nearest neighbor (ANN) algorithms (e.g., FAISS), followed by context-aware fine-grained re-ranking powered by large language models (LLMs). Distinct from prior approaches, our work is the first to deeply integrate LLMs into the vector search pipeline, leveraging customized prompt engineering and a structured evaluation framework to achieve precise semantic understanding of complex queries while maintaining millisecond-scale latency. Experimental results across multiple structured benchmarks demonstrate that our method improves accuracy by 32% on average over baseline vector-only search.
Traditional vector retrieval relies on approximate nearest neighbor (ANN) search, which often yields semantically redundant results and fails to meet the diversity and contextual richness requirements of applications such as retrieval-augmented generation (RAG) and multi-hop question answering. To address this, we propose a novel paradigm—“Semantic Compression and Graph-Enhanced Retrieval”—that introduces submodular optimization into vector retrieval for the first time. We formalize semantic compression to maximize information coverage while explicitly suppressing redundancy. Leveraging information-geometric similarity metrics and k-nearest neighbor (kNN) graphs, we construct a multi-hop semantic search framework, further augmented with knowledge graph integration for structured semantic querying. Our method supports hybrid indexing and significantly improves semantic diversity and coverage in high-dimensional embedding spaces, outperforming state-of-the-art ANN baselines. The implementation is open-sourced, advancing research toward semantics-centric vector search.
This study addresses the lack of systematic evaluation of hybrid search mechanisms that combine semantic retrieval with metadata filtering in existing vector databases. We propose a novel relevance metric, Global-Local Selectivity (GLS), construct MoReVec—the first benchmark dataset supporting filtered retrieval—and extend ANN-Benchmarks to enable unified evaluation of hybrid search performance. Through comprehensive experiments integrating diverse filtering strategies into FAISS, Milvus, and pgvector with IVFFlat and HNSW indexes, we demonstrate that engine-level algorithmic integration critically governs performance: Milvus achieves more stable recall via hybrid execution, pgvector’s optimizer often selects suboptimal query plans, and IVFFlat outperforms HNSW under low-selectivity queries. Our findings culminate in practical configuration guidelines that offer both theoretical insights and actionable recommendations for efficient hybrid search deployment.
Existing embedding-based retrieval systems employ fixed recall budgets, leading to insufficient recall for head queries and degraded precision for tail queries—rooted in the frequentist paradigm of loss functions. This paper proposes a query-aware probabilistic retrieval framework that abandons fixed truncation and instead models the query-specific distribution of candidate cosine similarities. By learning a query-conditioned cumulative distribution function (CDF), the framework dynamically determines similarity thresholds per query. For the first time, it jointly optimizes precision and recall for both head and tail queries within a unified probabilistic framework. Experiments across multiple industrial retrieval benchmarks demonstrate significant improvements in overall effectiveness. Ablation studies validate the efficacy of explicitly modeling head–tail disparities. The core innovation lies in elevating threshold selection from heuristic, static configuration to a principled, query-adaptive, probabilistic decision process.
This work proposes a novel approach to large-scale retrieval that circumvents the prohibitive cost of full reranking by constructing query and item embeddings derived from the outputs of a reranker. Specifically, it leverages relevance scores assigned by a heavyweight reranker over a set of support items to generate lightweight embeddings, thereby enabling the reranking model to directly guide embedding learning—a capability demonstrated here for the first time. Under mild conditions, the method is theoretically shown to approximate arbitrarily complex similarity functions. Through systematic investigation of support item selection strategies and integration with approximate nearest neighbor search, the approach significantly improves candidate set quality across multiple academic and industrial datasets while maintaining computational efficiency.
This work addresses the gap between benchmark-driven embedding model selection and real-world deployment constraints by introducing the first framework to evaluate embedding models within a complete retrieval pipeline. It systematically compares the end-to-end performance of T3EM’s commercial API against leading open-source models across diverse tasks—including retrieval, classification, clustering, and semantic similarity—as covered by the MTEB benchmark, while jointly accounting for latency, cost, task type, and deployment conditions. The study develops a comprehensive, end-to-end model selection guide encompassing embedding generation, indexing, search, and chunking strategies, revealing significant performance discrepancies that emerge only in full-system contexts. These insights provide practitioners with actionable, empirically grounded criteria for embedding model adoption in real-world applications.
This study addresses the lack of comprehensive, reproducible benchmarks for vector databases across retrieval quality, latency, throughput, and resource consumption. For the first time, it jointly evaluates seven systems—FAISS, Qdrant, Milvus, Weaviate, Chroma, pgvector, and LanceDB—across six datasets (including SIFT, GIST, MS MARCO, and GloVe) comprising over four million vectors, measuring fifteen performance and resource metrics. The evaluation reveals that FAISS achieves the highest single-node throughput (866 QPS), Weaviate delivers out-of-the-box recall exceeding 99%, Qdrant exhibits the lowest median latency (4.55 ms), and LanceDB constructs indexes fastest albeit with slightly lower retrieval quality. The authors open-source the complete benchmarking framework and propose practical guidelines for system selection.