Score
Designs, builds, and evaluates end-to-end retrieval-augmented generation (evidence-RAG) pipelines that ingest and chunk source material, create vector/semantic indexes and retrievers, select and rank supporting evidence, and fuse that evidence into generated outputs with provenance and retrieval traces exposed for editorial review or downstream UIs.
This study addresses the lack of systematic, up-to-date synthesis of retrieval-augmented generation (RAG) research amid rapid methodological diversification and evaluation fragmentation. Adopting the PRISMA 2020 framework, we systematically curated 128 highly cited papers (2020–2025) from ACM, IEEE, and other authoritative sources, introducing a dynamic citation threshold to mitigate temporal bias. Our analysis maps evolutionary trajectories across three dimensions: architectural design, benchmark datasets, and evaluation metrics—revealing critical methodological gaps, particularly in non-parametric memory augmentation and neural retrieval–generation co-adaptation. We construct a structured knowledge graph of RAG research and propose a prioritized roadmap that jointly optimizes robustness, interpretability, and generalization. The findings provide empirically grounded guidance for both foundational RAG theory development and practical deployment.
This work addresses the reliability challenges of Retrieval-Augmented Generation (RAG) in biomedical and clinical question answering by proposing a novel RAG framework that integrates hybrid retrieval, re-ranking, and claim-level evidence verification. The system leverages Amazon Bedrock for document processing and retrieval, enhances evidence relevance through Amazon Titan Text Embeddings V2, OpenSearch Serverless, and Cohere re-ranking, and employs a dedicated judgment model to rigorously verify each generated claim against source evidence. Evaluated on 25 biomedical queries, the framework successfully grounded all 200 extracted factual claims in their original evidence sources out of 500 retrieved passages, achieving 100.0% grounding accuracy and substantially improving the factual reliability and verifiability of generated responses.
This study investigates whether retrieval quality can serve as a reliable early indicator of information coverage in responses generated by Retrieval-Augmented Generation (RAG) systems. Through systematic experiments across three benchmarks—TREC NeuCLIR 2024, TREC RAG 2024, and WikiVideo—the authors evaluate 15 text-based and 10 multimodal retrieval systems using the Auto-ARGUE and MiRAGE assessment frameworks. The work provides the first empirical evidence of a strong correlation between coverage-oriented retrieval metrics and the informational coverage of generated outputs. Findings reveal that, at both topic and system levels, such metrics effectively predict RAG output coverage when retrieval and generation objectives are aligned, underscoring the critical role of goal consistency in optimizing RAG performance.
This work addresses the lack of an integrated Retrieval-Augmented Generation (RAG) development and evaluation toolkit in the R programming language, where existing solutions predominantly rely on the Python ecosystem. The authors propose ragR, the first native R framework enabling end-to-end RAG workflows—including document ingestion, vector storage, similarity-based retrieval, evidence synthesis, and structured question-answering logging—and faithfully reimplements the four core RAGAS evaluation metrics: context precision, context recall, faithfulness, and answer relevance. Experimental results demonstrate that ragR yields evaluation outcomes highly consistent with those produced by the original Python-based RAGAS implementation. This provides R users with a lightweight, reproducible, and self-contained environment for RAG research and education without requiring language switching.
This work addresses the bottleneck in RAG system evaluation—its heavy reliance on human-annotated ground-truth answers—by proposing RAGAs, a reference-free automated evaluation framework. Methodologically, it introduces a computable, three-dimensional metric suite covering retrieval relevance, context faithfulness, and generation quality, integrating BERTScore for semantic similarity, NLI-based models for factual consistency, and self-supervised prompting strategies. Its key contribution is the first end-to-end, multidimensional, reference-free evaluation paradigm, enabling quantitative, pipeline-level diagnostics of RAG systems. Experiments demonstrate strong agreement between automated metrics and human judgments (average Spearman ρ > 0.82) across multiple benchmarks, validating efficacy and robustness. The open-source RAGAs toolkit has been widely adopted in industry for iterative RAG system optimization.
This work addresses the limitations of existing approaches in multi-hop retrieval-augmented generation, which rely on fixed pipelines and lack dynamic control over evidence manipulation. The authors propose the first unified state-conditioned control framework, modeling multi-hop evidence acquisition as a sequence of atomic operations conditioned on the current reasoning state. A validity filtering layer constructs a feasible action set, from which a learnable controller adaptively selects the optimal operation. Integrating state-conditioned policy learning with the Qwen2.5-7B-Instruct model, the method is optimized end-to-end and achieves F1 scores of 0.5998, 0.5340, and 0.3061 on HotpotQA, 2WikiMultihopQA, and MuSiQue, respectively—significantly outperforming existing controllable baselines. Ablation studies confirm the critical contributions of both the learned controller and the sufficiency-based feedback mechanism.