evidence-aware rag

Designs and implements retrieval-augmented generation (RAG) systems that explicitly manage and incorporate external evidence: they build typed-query generation from failure signals, multi-stage (two-stage or staged) retrieval pipelines that fetch targeted evidence, and mechanisms to route, refresh, and evolve pools of evidence; they also stage retrieved evidence into the generation process in multiple steps to improve relevance and correctness.

evidence-awarerag

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.44
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Is Relevance Propagated from Retriever to Generator in RAG?

Feb 20, 2025
FT
Fangzheng Tian
🏛️ University of Glasgow

This study investigates whether topic relevance of retrieved documents in Retrieval-Augmented Generation (RAG) effectively improves downstream information retrieval performance. Moving beyond conventional “answer-containing” relevance, we propose and quantify a novel “topic-overlap” relevance dimension. Our methodology integrates the KILT benchmark, multi-scale context ablation experiments, standard retrievers (BM25, DPR), and utility attribution analysis. Results show that topic relevance exhibits only weak positive correlation with generation utility, and this correlation decays significantly as the number of retrieved documents (k) increases. While stronger retrieval models consistently improve overall RAG performance, they fail to translate topic relevance into utility gains linearly. The core contribution is the empirical and theoretical establishment of topic overlap as an independent relevance dimension, coupled with the discovery of a nonlinear attenuation law governing the relevance-to-utility transfer—highlighting diminishing returns in leveraging topic-relevant context as k grows.

Assessing relevance propagation in RAG systemsExploring topical overlap impact on RAG utilityInvestigating retrieval model effectiveness in RAG performance

Enhancing Retrieval-Augmented Generation: A Study of Best Practices

Jan 13, 2025
SL
Siran Li
🏛️ University of Tübingen

This study systematically investigates the impact mechanisms of individual components in Retrieval-Augmented Generation (RAG) systems on complex question answering and cross-domain tasks. Addressing key challenges—including low retrieval precision, weak contextual relevance, and poor multilingual adaptability—we propose three core innovations: (1) a Contrastive In-Context Learning (CICL) RAG paradigm to improve generation accuracy; (2) sentence-granularity focused retrieval (“Focus Mode”) combined with multi-granularity chunking to enhance retrieval relevance; and (3) a multilingual knowledge base integration framework that balances retrieval–generation efficiency. Through large-scale hyperparameter analysis, we quantitatively characterize the influence of critical factors—including language model scale, chunk size, and retrieval stride—on end-to-end performance. The findings yield a reproducible best-practice guideline for RAG system design and deployment, accompanied by open-sourced, fully implemented code.

cross-domain applicationperformance enhancementRAG system optimization

This study addresses the lack of standardized guidelines for deploying retrieval-augmented generation (RAG) systems in industrial healthcare settings. It presents the first systematic evaluation of RAG components on medical tasks, employing a modular architecture, ablation studies, and multi-task assessment to investigate the trade-offs between performance and efficiency. The work proposes a practical, evidence-based framework of best practices and identifies an optimal component configuration that significantly improves both answer accuracy and reasoning efficiency across three representative medical tasks. These findings provide empirical support and actionable technical guidance for the real-world deployment of medical RAG systems.

Best PracticesIndustrial ApplicationsLarge Language Models

Creating a Taxonomy for Retrieval Augmented Generation Applications

Aug 05, 2024
IN
Irina Nikishina
🏛️ University of Hamburg | University of Kassel

This work addresses the lack of a systematic taxonomy for retrieval-augmented generation (RAG) applications. We propose the first comprehensive, lifecycle-spanning classification framework for RAG applications. Methodologically, we introduce a novel four-stage iterative construction paradigm—comprising multi-round expert collaboration, systematic literature review, dimensional abstraction, and empirical validation—thereby filling a critical gap in classification research beyond the ACL community. The framework comprises five meta-dimensions and sixteen fine-grained dimensions, balancing structural clarity with extensibility. Empirical validation across education, healthcare, and legal domains demonstrates its effectiveness in supporting design decisions, technical evaluation, and cross-domain understanding of RAG applications. By providing a foundational taxonomic infrastructure, this work advances the engineering-oriented deployment and standardization of RAG systems.

Develop taxonomy for RAG applicationsEnhance understanding of RAG dimensionsFacilitate RAG adoption in various domains

Towards Understanding Retrieval Accuracy and Prompt Quality in RAG Systems

Nov 29, 2024
SZ
Shengming Zhao
🏛️ University of Alberta | The University of Tokyo | East China Normal University

The impact of key design decisions—RAG activation, retrieval granularity, and knowledge integration strategy—on RAG system performance remains poorly understood. Method: We conduct systematic ablation studies across three code/qa benchmarks and two state-of-the-art LLMs, quantitatively evaluating how document type, recall rate, document selection strategy, and prompt engineering jointly affect answer correctness and confidence via multi-dimensional analysis, cross-model/dataset comparison, and joint prompt-retrieval analysis. Contribution/Results: We identify precise interaction patterns and operational boundaries among these factors and propose nine actionable, empirically grounded guidelines for diagnosing and optimizing RAG failures. Our findings significantly improve RAG system stability, debuggability, and reliability, offering rigorous empirical evidence and a principled methodology to support the engineering deployment of LLM-augmented systems.

Analyzing key engineering trade-offs in RAG deployment decisionsDetermining optimal retrieval volume for different task typesEvaluating effective knowledge integration methods across tasks

Latest Papers

What's happening recently
View more

This work addresses the challenge of leveraging LaTeX source code—rich in structural and semantic information yet hindered by cross-references, custom macros, and unmarked content—for retrieval-augmented generation (RAG). The authors propose a systematic preprocessing pipeline that transforms raw LaTeX documents and their auxiliary files into structured Markdown and JSONL fragments through LaTeX parsing, macro expansion, reference resolution, and semantic annotation. This pipeline enables efficient indexing in vector databases and constitutes the first end-to-end method to convert native LaTeX into a knowledge format readily consumable by large language models. By preserving document structure, semantic labels, and authorial intent, the approach significantly enhances the accuracy and reliability of language models on mathematical and technical question-answering tasks.

AI-friendlyKnowledge SourceLaTeX

This work addresses the limitations of existing approaches in multi-hop retrieval-augmented generation, which rely on fixed pipelines and lack dynamic control over evidence manipulation. The authors propose the first unified state-conditioned control framework, modeling multi-hop evidence acquisition as a sequence of atomic operations conditioned on the current reasoning state. A validity filtering layer constructs a feasible action set, from which a learnable controller adaptively selects the optimal operation. Integrating state-conditioned policy learning with the Qwen2.5-7B-Instruct model, the method is optimized end-to-end and achieves F1 scores of 0.5998, 0.5340, and 0.3061 on HotpotQA, 2WikiMultihopQA, and MuSiQue, respectively—significantly outperforming existing controllable baselines. Ablation studies confirm the critical contributions of both the learned controller and the sufficiency-based feedback mechanism.

evidence controllearnable policymulti-hop retrieval-augmented generation

Hot Scholars

PM

Pekka Marttinen

Aalto University
Statistical machine learningComputational biology
HX

Hongxia Xu

Zhejiang University
AI4ScienceNanomedicineMedical imaging
FW

Furu Wei

Distinguished Scientist, Microsoft Research
Natural Language ProcessingArtificial IntelligenceGeneral AIGenerative AI