Score
Designs, builds, and integrates pipelines that condition generative models on retrieved external evidence by combining retrievers (lexical, dense/latent, hybrid, exemplar-guided, section-aware, multimodal, on-device or offline) with prompt- or model-level augmentation to produce grounded outputs. Analyzes and evaluates indexing and retrieval components, retrieval-to-generation interfaces, grounding and verification modules, and metrics for faithfulness, hallucination reduction, pass@k and label-space narrowing.
This study addresses the lack of systematic, up-to-date synthesis of retrieval-augmented generation (RAG) research amid rapid methodological diversification and evaluation fragmentation. Adopting the PRISMA 2020 framework, we systematically curated 128 highly cited papers (2020–2025) from ACM, IEEE, and other authoritative sources, introducing a dynamic citation threshold to mitigate temporal bias. Our analysis maps evolutionary trajectories across three dimensions: architectural design, benchmark datasets, and evaluation metrics—revealing critical methodological gaps, particularly in non-parametric memory augmentation and neural retrieval–generation co-adaptation. We construct a structured knowledge graph of RAG research and propose a prioritized roadmap that jointly optimizes robustness, interpretability, and generalization. The findings provide empirically grounded guidance for both foundational RAG theory development and practical deployment.
Traditional retrievers rely on surface-level matching and struggle to capture complex user intent, resulting in a semantic gap between queries and documents. This work proposes the Generative Embedding Model (GEM), which uniquely integrates explicit reasoning into the embedding architecture: it first employs a large language model to perform intent and relevance reasoning over the query, then appends embedding tokens that encode this enhanced contextual representation for retrieval. The unified framework enables test-time computational scaling through prompting and significantly outperforms non-reasoning baselines on both reasoning-intensive and instruction-following retrieval tasks. Notably, GEM achieves performance comparable to that of substantially larger models while operating at a smaller scale.
To address the limited output quality of Large Vision-Language Models (LVLMs) in multimodal understanding and generation, this paper proposes UniRAG—a plug-and-play retrieval-augmented reasoning framework. At its core, UniRAG leverages a vision-language retriever (e.g., UniIR) to dynamically retrieve relevant multimodal examples during inference and inject them as few-shot demonstrations into the prompt—requiring neither model fine-tuning nor architectural modification. This work presents the first systematic empirical validation that retrieval augmentation significantly improves LVLM performance in common-entity scenarios, demonstrating strong cross-model generalizability. On the MSCOCO benchmark, UniRAG consistently enhances image captioning quality across diverse state-of-the-art LVLMs, including GPT-4o, Gemini-Pro, LLaVA, LaVIT, and Emu2. The implementation is publicly available.
To address hallucination—i.e., the generation of factually unsupported content—in large language model (LLM)-based chatbots, this paper proposes a training-free, plug-and-play post-hoc method. Given a user query, the method first generates an initial response; then retrieves supporting documents via BM25 or embedding-based search; next employs a natural language inference (NLI) model to verify the factual grounding of each claim against the retrieved evidence; and finally iteratively regenerates the response until all claims are fully attributable to cited sources. This work introduces the novel paradigm of “training-agnostic posterior citation augmentation,” enabling fully verifiable and traceable response generation without modifying or fine-tuning the underlying LLM. The method is compatible with any black-box LLM out-of-the-box. Evaluated on three established hallucination benchmarks, it substantially outperforms state-of-the-art approaches, achieving over 8% absolute improvement in both F1 (hallucination detection) and BLEU (response regeneration) scores.
Generative retrieval—where large language models (LLMs) autoregressively generate document identifiers—lacks a clear understanding of how model size, training data volume, and inference compute jointly scale. Method: We conduct the first systematic study of this triadic scaling relationship, introducing a continuous evaluation metric that integrates contrastive entropy and generative loss to enable robust, architecture-agnostic comparisons. Using a unified framework across LLaMA (decoder-only) and T5 (encoder-decoder), we combine n-gram analysis with large-scale ablation experiments. Contributions/Results: All three resources—model scale, data volume, and inference compute—exhibit strong positive correlations with retrieval performance. LLaMA consistently outperforms T5 across multiple configurations. Crucially, n-gram modeling adheres strictly to power-law scaling behavior, offering a novel, interpretable foundation for generative retrieval. Our findings establish principled guidelines for resource-aware model design and deployment in generative retrieval systems.
This work addresses the limitations of current dense retrieval methods, which either underutilize the reasoning capabilities of large language models (LLMs) by treating them as static encoders or suffer from high latency due to autoregressive generation in explicit reasoning. To bridge this gap, we propose LaSER, a novel framework that internalizes explicit reasoning paths into the retriever’s latent space through trajectory alignment, enabling efficient implicit reasoning without intermediate text generation. LaSER employs a dual-view training mechanism—explicit and implicit—on a shared LLM backbone, enhanced by multi-granularity alignment and self-distillation to jointly optimize reasoning depth and retrieval efficiency. Extensive experiments demonstrate that LaSER significantly outperforms existing approaches on both in-domain and cross-domain reasoning-intensive retrieval benchmarks, exhibiting strong robustness and effectiveness across varying model scales.
This work addresses the lack of systematic understanding of the differences and complementarities among diverse retrievers in current Retrieval-Augmented Generation (RAG) systems, which hinders effective retriever selection and integration. To this end, we propose MIGRASCOPE, a novel framework that introduces mutual information and statistical estimation theory into RAG evaluation, enabling quantitative assessment of retrievers in terms of retrieval quality, redundancy, synergy, and marginal contribution. Leveraging this framework, we uncover complementary relationships among mainstream retrievers and design efficient integration strategies. Experimental results demonstrate that carefully orchestrated multi-retriever systems significantly outperform the best single retriever, offering both theoretical grounding and practical guidance for building more effective and robust RAG systems.
This work addresses the limitations of existing retrieval-augmented generation (RAG) systems, which rely on explicit natural language queries and employ disjoint retriever and generator components, thereby failing to fully exploit the representational capacity of large language models. To overcome this, the authors propose the LAnR framework, which for the first time enables end-to-end joint encoding, retrieval, and generation within a unified latent space. LAnR generates dense retrieval vectors directly from the hidden states of a [PRED] token, eliminating the need for a separate retrieval module and explicit queries. Additionally, it introduces a lightweight MLP-based control head that adaptively assesses retrieval sufficiency via answer entropy, allowing early termination when further retrieval is unnecessary. Evaluated on six question-answering benchmarks, LAnR outperforms current RAG approaches while reducing retrieval calls, significantly enhancing both inference efficiency and factual accuracy.
This work addresses the challenges of inefficient retrieval and source-unverified generation in generative AI, which often lead to hallucinations. To mitigate these issues, the authors propose the MPR-CiteG framework, which integrates a Multi-Path Retriever (MPR) to efficiently gather diverse, relevant information and a Citation-Anchored Generation module (CiteG) that ensures every generated statement is grounded in accurate, traceable sources. This framework represents the first synergistic integration of multi-path retrieval with citation-anchored generation. Evaluated on the ScienceON AI Challenge dataset, MPR-CiteG significantly reduces hallucination rates while enhancing answer accuracy and credibility, securing second place in the competition and demonstrating marked improvements in factual consistency and reliability of large language model outputs.