Score
Designs and implements question-answering systems that integrate document retrieval with answer generation or extraction, building and evaluating retrieval components, rankers, context selectors, and grounding mechanisms. Produces pipelines that fetch domain-relevant documents, rank and select supporting contexts, and generate answers grounded in retrieved evidence with citations to reduce hallucinations.
To address core limitations of large language models (LLMs)—including hallucination, knowledge obsolescence, and poor domain adaptability—this work systematically advances the Retrieval-Augmented Structured (RAS) generation paradigm. We propose a multi-granularity knowledge acquisition mechanism integrating sparse, dense, and hybrid retrieval, coupled with text structuralization, taxonomy construction, knowledge embedding, and prompt-driven reasoning to enable efficient external knowledge retrieval, semantic alignment, and controllable integration. Crucially, we deeply embed structured modeling into the augmentation pipeline, enhancing factual accuracy, temporal freshness, and domain-specific competence of generated outputs. Our contributions include: (1) a unified methodological framework for RAS generation; (2) principled pathways toward multimodal, cross-lingual, and interactive augmented generation; and (3) empirically validated improvements in reliability and specialization across diverse domains. This work establishes foundational design principles and future research directions for next-generation RAS systems.
To address challenges in long-document multi-hop question answering—including difficulty in cross-paragraph reasoning, weak contextual understanding, and answer inconsistency—this paper proposes a retrieval-augmented generation (RAG) framework based on LLaMA-3. Methodologically: (1) it introduces a joint optimization objective that simultaneously minimizes retrieval likelihood loss and generation cross-entropy loss; (2) it designs a hierarchical multi-hop retrieval–generation coordination mechanism integrating a DPR variant, a multi-hop graph reasoning module, and a context-aware fusion encoder; and (3) it employs end-to-end joint training. The framework achieves state-of-the-art performance on HotpotQA, MuSiQue, and DocQA—improving multi-hop accuracy by 7.2% and answer faithfulness by 11.5% over baseline RAG and pure generative methods. Its core contribution lies in the first unified modeling of retrieval and generation objectives, coupled with deep synergy between multi-hop logical reasoning and contextual representation learning.
This study addresses a critical oversight in current retrieval-augmented generation (RAG) systems: their reliance on human-oriented document representations, which neglect the distinct representational needs of large language models as content consumers. Under fixed retrieval results, the authors systematically evaluate the impact of 14 document representation strategies—including selection, summarization, and rewriting—on question-answering accuracy across four generative models. Introducing answer retention rate as a novel metric to assess whether transformed documents preserve the correct answer, controlled experiments reveal for the first time that answer retention is the primary driver of generation accuracy, challenging prior assumptions that attributed performance gains to specific representational mechanisms. Notably, when answer retention is high, variations in wording, structure, length, or query dependence exert minimal influence on accuracy, underscoring that preserving answer information outweighs representational form.
Existing RAG systems struggle with “how-to” questions due to logical fragmentation and step disconnection, primarily caused by conventional fixed-size chunking that disrupts the semantic coherence of procedural knowledge. To address this, we propose Thread—a novel data organization paradigm grounded in logically cohesive, semantically self-contained units. Leveraging large language models, documents are parsed and reconstructed into loosely coupled reasoning threads, enabling fine-grained, process-aware knowledge representation. Our method integrates logic-driven document segmentation, cross-format adaptive indexing, and retrieval-augmented generation. Experiments across open-domain and industrial benchmarks demonstrate a 21–33% improvement in “how-to” question resolution success rate, a 75% reduction in candidate knowledge volume, and substantial gains in retrieval efficiency and answer executability.
Large language models (LLMs) suffer from hallucination in question answering due to conflicts between parametric knowledge and input context. Method: This paper proposes a novel paradigm that explicitly models entity–context alignment relationships. It is the first to systematically reveal the critical impact of entity–context alignment quality during training on inference faithfulness, and introduces an intervenable knowledge conflict mitigation mechanism—incorporating entity-aware attention constraints and context-alignment regularization loss into a seq-to-seq framework. Contribution/Results: The method significantly reduces hallucination rates across multiple document-level QA benchmarks (average reduction of 32.7%), while simultaneously improving answer faithfulness and context utilization. It establishes an interpretable and controllable technical pathway for trustworthy generative QA, enabling explicit intervention in knowledge grounding and alignment.
This work addresses the challenges of ambiguous citation provenance and content redundancy commonly encountered in existing retrieval-augmented generation (RAG) systems during information integration. The authors propose a knowledge base construction approach grounded in Q&A nuggets, which leverages explicit question-answer semantics to guide information extraction, selection, and generation while preserving source attribution throughout the pipeline. Departing from conventional fuzzy clustering abstractions, the method employs interpretable Q&A fragments as structured intermediate representations, enabling end-to-end traceable reasoning and generation. Experimental results on the TREC NeuCLIR 2024 dataset demonstrate that the proposed approach significantly outperforms the state-of-the-art nugget-based RAG system, Ginger, in terms of nugget recall, density, and citation accuracy.
This work addresses the lack of a unified analytical framework for retrieval and reasoning pipelines in multi-hop question answering, which has hindered systematic comparison across methods. We propose, for the first time, a four-axis design framework that treats the execution process as the fundamental unit of analysis, encompassing execution plans, index structures, control strategies, and termination criteria. Through a comprehensive literature review and ablation studies, we structurally map prominent approaches—including RAG and agent-based systems—onto this framework using benchmarks such as HotpotQA, revealing consistent trade-offs among effectiveness, efficiency, and evidence faithfulness. Our framework systematically organizes existing design choices, identifies reproducible empirical trends, and highlights key challenges, including structure-aware planning and transferable control strategies.
This work addresses the susceptibility of large language models to factual hallucinations in complex question answering. To mitigate this issue, the authors propose a lightweight graph-augmented retrieval-augmented generation system that integrates vector retrieval with graph-based query tools, enabling efficient multi-hop reasoning over a Wikipedia subset. By incorporating a concise graph schema and a dedicated toolset, the approach effectively curbs hallucinations and enhances fine-grained factual accuracy without substantially increasing computational overhead. Experimental results demonstrate a 50% reduction in hallucinated responses compared to baseline methods, along with significant improvements in both precision and recall for factual correctness. The method achieves state-of-the-art performance on the MoNaCo benchmark in terms of fine-grained truthfulness.
This work addresses the common oversight of internal document structure in existing document-based question answering systems, which often leads to incoherent retrieval and generation. The authors propose SF-Re2G, a novel framework that systematically integrates hierarchical document structure throughout the entire pipeline—retrieval, reranking, and generation. Specifically, structure-aware contrastive learning is introduced during retrieval to enhance paragraph representations within the same section; during reranking, a subgraph based on structural proximity is constructed to aggregate contextual information from neighboring passages; and finally, this subgraph context guides answer generation. Experimental results demonstrate that SF-Re2G significantly improves both retrieval and generation performance on Chinese and English document dialogue benchmarks, confirming the effectiveness and generalizability of leveraging structural information.
This work addresses the challenge that complex queries often admit multiple valid answers, which existing retrieval methods struggle to comprehensively cover. To this end, the authors propose the RVR framework, which employs an iterative closed-loop mechanism of “retrieve–verify–retrieve” to dynamically enrich the original query with verified documents, thereby enabling effective query expansion and answer discovery—without relying on complex agent-based systems. The approach requires only off-the-shelf or minimally fine-tuned retrievers and verifiers, yet achieves substantial gains in multi-answer recall. On the QAMPARI dataset, it yields a relative improvement of over 10% (3% absolute) in full-recall performance and consistently outperforms strong baselines across diverse domains, including QUEST and WebQuestionsSP.