🤖 AI Summary
This study addresses the performance bottleneck encountered by dense retrievers in oblique retrieval, where latent attributes lack explicit surface realizations. To overcome this limitation, this work proposes a single-vector dense retrieval framework that performs knowledge distillation from k-nearest neighbor (kNN) graphs using a frozen author encoder. By leveraging synthetic queries and cross-model supervision, the approach enhances implicit relevance recognition while effectively transferring stylistic and topical inductive biases to the student model, further integrated with contrastive learning for efficient representation alignment. Experimental results demonstrate that the proposed method significantly outperforms GPT-5.2 agent and Gemini-2-Embedding baselines in NDCG@10 across multiple benchmarks, thereby validating the efficacy of knowledge distillation in complex retrieval tasks.
📝 Abstract
Oblique retrieval, as exemplified by OBLIQ-Bench, asks a retriever to find documents whose relevance is determined by a latent attribute (an implicit stance, an analogous reasoning technique, an authorial fingerprint, or a vague tip-of-the-tongue recollection) that has little or no surface expression in the document. State-of-the-art dense encoders and agentic search pipelines built around frontier language models exhibit a large first-stage bottleneck on these tasks, while the same language models reliably verify relevance when shown candidates. We address this with OBLIQ-IR, a single-vector dense retriever whose training mixture combines per-mechanism synthetic queries with a new form of cross-model supervision: kNN-graph distillation from a frozen authorship encoder, which transfers a style-versus-topic inductive bias into the student. A 3B retriever fine-tuned reaches 0.211 NDCG@10 on Writing-Style, 0.171 on Math, 0.177 on Twitter, and 0.281 on Congress, improving over the GPT-5.2 Multi-Hop Agent by \xr{0.010 to 0.150} NDCG@10 and over Gemini-2-Embedding by 0.027 to 0.222 NDCG@10 on every reported task. The code, data and checkpoints are available https://github.com/DataScienceUIBK/obliq-ir