🤖 AI Summary
This study addresses the high learning difficulty of query encoders in vector retrieval, which causes practical recall rates to fall substantially below their geometric capacity upper bounds. By analyzing maximum recall under a frozen document index, this work proposes a computational hardness theory grounded in statistical queries. It proves that even with a perfect encoder, surpassing random baselines requires exponentially many samples. This theoretical derivation is complemented by empirical validation through minimum embedding dimension analysis and single-layer ReLU network constructions. The contributions confirm the existence of unrealized geometric capacity in mainstream benchmarks and establish that the learnability of query encoders constitutes the critical bottleneck constraining the performance of embedding-based retrieval systems.
📝 Abstract
Efficient vector retrieval requires both a corpus geometry that supports retrieving the right documents through vector similarity, and a query encoder that can embed queries near their desired documents in the embedding space. Recent work has studied geometric capacity through the lens of the minimum embedding dimension needed to realize all top-$k$ answer sets of $n$ documents. We study a different notion of geometric capacity--the maximum recall achievable for a frozen document index--and explore whether learned query encoders can reach this ceiling. On several real-world retrieval benchmarks, we show that retrieval quality of single-vector query encoders often lies far below what the document indices can support.
Motivated by this observation, we give theoretical evidence that learning query encoders can be computationally hard. In particular, we construct a retrieval task that (1) admits a query encoder with perfect recall which is representable by a small one-hidden-layer ReLU network, but (2) any statistical-query learner (a class capturing learners that access training data through aggregate statistics) provably requires exponentially many statistical queries to achieve non-trivial recall advantage over the random baseline $k/n$. Taken together, our results suggest substantial unrealized geometric capacity in retrieval benchmarks and establish query encoder learnability as a possible barrier in embedding-based retrieval.