retrieval-augmented qa

Designs and implements question-answering systems that integrate document retrieval with answer generation or extraction, building and evaluating retrieval components, rankers, context selectors, and grounding mechanisms. Produces pipelines that fetch domain-relevant documents, rank and select supporting contexts, and generate answers grounded in retrieved evidence with citations to reduce hallucinations.

retrieval-augmentedqa

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Enhancing Document-Level Question Answering via Multi-Hop Retrieval-Augmented Generation with LLaMA 3

Jun 19, 2025
XH
Xinyue Huang
🏛️ Cornell University | University of Southern California

To address challenges in long-document multi-hop question answering—including difficulty in cross-paragraph reasoning, weak contextual understanding, and answer inconsistency—this paper proposes a retrieval-augmented generation (RAG) framework based on LLaMA-3. Methodologically: (1) it introduces a joint optimization objective that simultaneously minimizes retrieval likelihood loss and generation cross-entropy loss; (2) it designs a hierarchical multi-hop retrieval–generation coordination mechanism integrating a DPR variant, a multi-hop graph reasoning module, and a context-aware fusion encoder; and (3) it employs end-to-end joint training. The framework achieves state-of-the-art performance on HotpotQA, MuSiQue, and DocQA—improving multi-hop accuracy by 7.2% and answer faithfulness by 11.5% over baseline RAG and pure generative methods. Its core contribution lies in the first unified modeling of retrieval and generation objectives, coupled with deep synergy between multi-hop logical reasoning and contextual representation learning.

Enhancing document-level QA via multi-hop RAGImproving multi-hop reasoning in lengthy documentsOptimizing retrieval-generation synergy for accurate answers

This study addresses a critical oversight in current retrieval-augmented generation (RAG) systems: their reliance on human-oriented document representations, which neglect the distinct representational needs of large language models as content consumers. Under fixed retrieval results, the authors systematically evaluate the impact of 14 document representation strategies—including selection, summarization, and rewriting—on question-answering accuracy across four generative models. Introducing answer retention rate as a novel metric to assess whether transformed documents preserve the correct answer, controlled experiments reveal for the first time that answer retention is the primary driver of generation accuracy, challenging prior assumptions that attributed performance gains to specific representational mechanisms. Notably, when answer retention is high, variations in wording, structure, length, or query dependence exert minimal influence on accuracy, underscoring that preserving answer information outweighs representational form.

answer retentiondocument transformationlarge language model

Thread: A Logic-Based Data Organization Paradigm for How-To Question Answering with Retrieval Augmented Generation

Jun 19, 2024
KA
Kaikai An
🏛️ Peking University | Microsoft | Nanjing University

Existing RAG systems struggle with “how-to” questions due to logical fragmentation and step disconnection, primarily caused by conventional fixed-size chunking that disrupts the semantic coherence of procedural knowledge. To address this, we propose Thread—a novel data organization paradigm grounded in logically cohesive, semantically self-contained units. Leveraging large language models, documents are parsed and reconstructed into loosely coupled reasoning threads, enabling fine-grained, process-aware knowledge representation. Our method integrates logic-driven document segmentation, cross-format adaptive indexing, and retrieval-augmented generation. Experiments across open-domain and industrial benchmarks demonstrate a 21–33% improvement in “how-to” question resolution success rate, a 75% reduction in candidate knowledge volume, and substantial gains in retrieval efficiency and answer executability.

Addressing logical coherence loss in how-to question answeringEnhancing retrieval efficiency for complex procedural informationImproving step-by-step reasoning for dynamic decision-making processes

Mitigating Knowledge Conflicts in Language Model-Driven Question Answering

Nov 18, 2024
HC
Han Cao
🏛️ University of California, San Diego | Washington University in St. Louis

Large language models (LLMs) suffer from hallucination in question answering due to conflicts between parametric knowledge and input context. Method: This paper proposes a novel paradigm that explicitly models entity–context alignment relationships. It is the first to systematically reveal the critical impact of entity–context alignment quality during training on inference faithfulness, and introduces an intervenable knowledge conflict mitigation mechanism—incorporating entity-aware attention constraints and context-alignment regularization loss into a seq-to-seq framework. Contribution/Results: The method significantly reduces hallucination rates across multiple document-level QA benchmarks (average reduction of 32.7%), while simultaneously improving answer faithfulness and context utilization. It establishes an interpretable and controllable technical pathway for trustworthy generative QA, enabling explicit intervention in knowledge grounding and alignment.

Knowledge ContradictionLanguage ModelsQuestion Answering

This work addresses the challenges of ambiguous citation provenance and content redundancy commonly encountered in existing retrieval-augmented generation (RAG) systems during information integration. The authors propose a knowledge base construction approach grounded in Q&A nuggets, which leverages explicit question-answer semantics to guide information extraction, selection, and generation while preserving source attribution throughout the pipeline. Departing from conventional fuzzy clustering abstractions, the method employs interpretable Q&A fragments as structured intermediate representations, enabling end-to-end traceable reasoning and generation. Experimental results on the TREC NeuCLIR 2024 dataset demonstrate that the proposed approach significantly outperforms the state-of-the-art nugget-based RAG system, Ginger, in terms of nugget recall, density, and citation accuracy.

citation provenanceinformation redundancyinterpretable generation

Latest Papers

What's happening recently
View more

This work addresses the lack of a unified analytical framework for retrieval and reasoning pipelines in multi-hop question answering, which has hindered systematic comparison across methods. We propose, for the first time, a four-axis design framework that treats the execution process as the fundamental unit of analysis, encompassing execution plans, index structures, control strategies, and termination criteria. Through a comprehensive literature review and ablation studies, we structurally map prominent approaches—including RAG and agent-based systems—onto this framework using benchmarks such as HotpotQA, revealing consistent trade-offs among effectiveness, efficiency, and evidence faithfulness. Our framework systematically organizes existing design choices, identifies reproducible empirical trends, and highlights key challenges, including structure-aware planning and transferable control strategies.

evidence faithfulnessexecution proceduremulti-hop question answering

This work addresses the susceptibility of large language models to factual hallucinations in complex question answering. To mitigate this issue, the authors propose a lightweight graph-augmented retrieval-augmented generation system that integrates vector retrieval with graph-based query tools, enabling efficient multi-hop reasoning over a Wikipedia subset. By incorporating a concise graph schema and a dedicated toolset, the approach effectively curbs hallucinations and enhances fine-grained factual accuracy without substantially increasing computational overhead. Experimental results demonstrate a 50% reduction in hallucinated responses compared to baseline methods, along with significant improvements in both precision and recall for factual correctness. The method achieves state-of-the-art performance on the MoNaCo benchmark in terms of fine-grained truthfulness.

factual correctnesshallucinationlarge language models

This work addresses the common oversight of internal document structure in existing document-based question answering systems, which often leads to incoherent retrieval and generation. The authors propose SF-Re2G, a novel framework that systematically integrates hierarchical document structure throughout the entire pipeline—retrieval, reranking, and generation. Specifically, structure-aware contrastive learning is introduced during retrieval to enhance paragraph representations within the same section; during reranking, a subgraph based on structural proximity is constructed to aggregate contextual information from neighboring passages; and finally, this subgraph context guides answer generation. Experimental results demonstrate that SF-Re2G significantly improves both retrieval and generation performance on Chinese and English document dialogue benchmarks, confirming the effectiveness and generalizability of leveraging structural information.

contextdocument structuredocument-grounded dialogue

This work addresses the challenge that complex queries often admit multiple valid answers, which existing retrieval methods struggle to comprehensively cover. To this end, the authors propose the RVR framework, which employs an iterative closed-loop mechanism of “retrieve–verify–retrieve” to dynamically enrich the original query with verified documents, thereby enabling effective query expansion and answer discovery—without relying on complex agent-based systems. The approach requires only off-the-shelf or minimally fine-tuned retrievers and verifiers, yet achieves substantial gains in multi-answer recall. On the QAMPARI dataset, it yields a relative improvement of over 10% (3% absolute) in full-recall performance and consistently outperforms strong baselines across diverse domains, including QUEST and WebQuestionsSP.

answer coveragecomplete recallcomprehensive question answering

Hot Scholars

RK

Ranjay Krishna

University of Washington, Allen Institute for AI
Computer VisionNatural Language ProcessingMachine LearningHuman Computer Interaction
JG

Jiuxiang Gu

Adobe Research
Computer VisionNatural Language ProcessingMachine Learning
SO

Sarah Ostadabbas

Electrical & Computer Engineering, Northeastern University
Computer VisionMachine LearningArtificial IntelligenceAugmented Cognition with Medical
DZ

Ding Zhao

Carnegie Mellon University
Trustworthy AIAI safetyreinforcement learningautonomous vehicles
SR

Sailaja Rajanala

Monash University Malaysia
NLPCausalityRepresentation Learning