🤖 AI Summary
This study addresses the fundamental disconnect between long-context reasoning and retrieval-augmented generation regarding evidence selection granularity. To bridge this gap, we propose a model-native evidence selection framework that leverages the internal representations of frozen large language models to unify corpus retrieval and long-text reasoning. By employing representation-based encoding and query generation techniques integrated with dense and hybrid retrieval architectures, the framework achieves cross-granularity evidence selection through the introduction of minimal additional parameters, without modifying the backbone model. Empirical evaluations demonstrate that our approach substantially reduces computational overhead in terms of FLOPs while improving recall to 73.2% on HotpotQA and yielding significant accuracy gains across long-context tasks.
📝 Abstract
Long-context inference and Retrieval-Augmented Generation (RAG) handle evidence selection at vastly different scales, from a single long prompt to an entire corpus. We ask whether a single model-internal mechanism can select evidence across this range. We introduce UNifying REtrieval And Long-Context with a Single Model (UNREAL), a model-native evidence selection framework to span corpus retrieval and long-context inference. UNREAL encodes chunks and derives retrieval queries directly from the frozen LLM's internal representations. It adds fewer than 500K trainable parameters and leaves the backbone unchanged. On a 3B-token, 21M-chunk Wikipedia index, all four dense and hybrid UNREAL backbones outperform state-of-the-art retriever-reranker systems. The best model raises recall from 49.1% to 73.2% on HotpotQA, from 31.7% to 60.1% on 2WikiMultiHopQA, and from 8.8% to 14.4% on MuSiQue. Applied to long-context tasks, the same selection mechanism removes distractors before generation, raising NoLiMa accuracy from 1.0% to 24.83% at its maximum context length of 128K tokens, and LV-Eval's F1 score from 49.97% to 54.66% at 256K. UNREAL also reduces FLOPs and time-to-first-token relative to full-context inference from roughly 32K tokens onward, with larger gains as context grows. Together, these results establish model-internal evidence selection as a common foundation for corpus retrieval and evidence-sparse long-context inference.