Latent Abstraction for Retrieval-Augmented Generation

📅 2026-04-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing retrieval-augmented generation (RAG) systems, which rely on explicit natural language queries and employ disjoint retriever and generator components, thereby failing to fully exploit the representational capacity of large language models. To overcome this, the authors propose the LAnR framework, which for the first time enables end-to-end joint encoding, retrieval, and generation within a unified latent space. LAnR generates dense retrieval vectors directly from the hidden states of a [PRED] token, eliminating the need for a separate retrieval module and explicit queries. Additionally, it introduces a lightweight MLP-based control head that adaptively assesses retrieval sufficiency via answer entropy, allowing early termination when further retrieval is unnecessary. Evaluated on six question-answering benchmarks, LAnR outperforms current RAG approaches while reducing retrieval calls, significantly enhancing both inference efficiency and factual accuracy.

Technology Category

Natural Language Processing: Question AnsweringMachine Learning: Large Multimodal Models (LMMs)Search and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Retrieval-Augmented Generation (RAG) has become a standard approach for enhancing large language models (LLMs) with external knowledge, mitigating hallucinations, and improving factuality. However, existing systems rely on generating natural language queries at each hop and maintaining a strict architectural separation between retriever and generator, preventing them from leveraging the full representational capacity of the LLM. We propose \textbf{LAnR} (Latent Abstraction for RAG), a unified framework in which a single LLM jointly performs encoding, retrieval, and generation entirely within its own latent space. Rather than generating textual queries, LAnR produces dense retrieval vectors from the hidden states of a designated \texttt{[PRED]} token and uses them to match against encoded document representations from the same model. Furthermore, LAnR adaptively decides when sufficient evidence has been retrieved using a lightweight MLP control head over those same hidden states, eliminating both the separate retriever and explicit token-level stopping reasoning. This design is motivated by our empirical observation that answer token entropy reliably signals retrieval sufficiency. Extensive experiments on six QA benchmarks spanning single-hop and multi-hop settings demonstrate that LAnR outperforms existing RAG methods, while achieving improved inference efficiency through reduced number of retrieval calls and tighter model integration.
Problem

Research questions and friction points this paper is trying to address.

Retrieval-Augmented Generation
large language models
latent space
retriever-generator separation
query generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Retrieval-Augmented Generation
Latent Space Integration
Dense Retrieval Vectors
Adaptive Retrieval Control
Unified LLM Framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Ha Lan N. T
Center for AI Research, VinUniversity, Vietnam
M
Minh-Anh Nguyen
Center for AI Research, VinUniversity, Vietnam
D
Dung D. Le
Center for AI Research, VinUniversity, Vietnam