CeQe: Grounding Lexical Retrieval in Semantic Evidence

📅 2026-08-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the issue of relevant document omission in lexical retrieval methods like BM25, which stems from the semantic vocabulary gap. To mitigate this, the authors propose Cross-Encoder Query Expansion (CE-QE), a method that leverages token-level relevance attributions from cross-encoders applied to semantically retrieved documents to extract key terms for query expansion. By incorporating these terms into the original query without modifying the index, CE-QE enhances lexical recall while avoiding the self-reinforcing bias of traditional pseudo-relevance feedback and the hallucination risks of generative expansion. Integrated with BM25, semantic retrieval, and a novel score fusion strategy (SESF), the system achieves substantial performance gains across seven BEIR datasets—for instance, Recall@100 on NQ improves from 0.32 to 0.47. SESF further outperforms cross-encoder fusion by 2.5% in Recall@100 and surpasses SPLADEv2 and ColBERTv2 by 5.3% and 4.6% in nDCG@10, respectively.
📝 Abstract
Lexical retrieval (BM25) captures exact keyword matches and weights terms by corpus-wide significance, but it is blind to the semantic vocabulary gap: when a relevant document phrases an answer differently from the query, BM25 never retrieves it, and no amount of downstream reranking or fusion can recover a document that was never in the candidate set. We present Cross-Encoder Query Expansion (CE-QE), which reads the per-token relevance attributions of a cross-encoder applied to top semantic search results, selects the terms the cross-encoder treats as decisive, and appends them to the BM25 query. Unlike classical pseudo-relevance feedback, which reuses BM25's own (possibly wrong) top results, CE-QE seeds expansion from the semantic retriever's results, avoiding self-reinforcing query drift. Unlike recent generative query expansion (HyDE, Query2doc), which prompts a large language model to hallucinate text from its parametric knowledge, every CE-QE expansion term is copied verbatim from a retrieved passage, so it cannot introduce vocabulary the corpus does not contain, and its only added cost is attribution extraction on a cross-encoder a hybrid pipeline already runs for reranking. On seven BEIR datasets, CE-QE improves lexical recall substantially where query and answer vocabulary diverge (e.g., NQ Recall@100 from 0.32 to 0.47), and its score-fusion variant (SESF) beats cross-encoder score fusion by 2.5% on Recall@100 and beats SPLADEv2 and ColBERTv2 by 5.3% and 4.6% on nDCG@10, while leaving the underlying BM25 index completely unmodified.
Problem

Research questions and friction points this paper is trying to address.

lexical retrieval
semantic vocabulary gap
query expansion
BM25
retrieval recall
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Encoder Query Expansion
lexical retrieval
semantic vocabulary gap
query expansion
relevance attribution
🔎 Similar Papers
No similar papers found.