Score
Designs and builds retrieval-augmented generation systems whose retrieval, re-ranking, fusion, and generation components are explicitly conditioned on a specified perspective or stakeholder viewpoint to produce perspective-aware outputs. This includes engineering multi-corpus evidence retrieval and fusion, perspective-specific filtering or re-ranking, and constrained-retrieval mechanisms and analyses to align evidence across perspectives and mitigate hallucination risk.
Hallucination remains a critical reliability challenge in the practical deployment of large language models (LLMs). Method: This work introduces, for the first time, a dual-dimensional taxonomy of hallucinations—distinguishing *knowledge-based* and *logic-based* types—and proposes a unified framework integrating retrieval-augmented generation (RAG), chain-of-thought (CoT) reinforcement, and agent-based system orchestration to systematically mitigate them. We analyze the intrinsic mechanisms by which each component suppresses distinct hallucination categories. Contribution/Results: Through rigorous empirical evaluation on standardized benchmarks, we systematically characterize the suppression pathways of each technique across hallucination types. The study delivers a reusable, modular paradigm for enhancing LLM reliability and a standardized evaluation framework. Our approach significantly improves both factual accuracy and operational feasibility—bridging the gap between theoretical robustness and real-world deployment.
Multi-source retrieval-augmented generation (RAG) suffers from data sparsity and cross-source information conflicts, significantly exacerbating hallucination. To address these challenges, we propose a knowledge-guided multi-source graph augmentation framework. First, we construct a multi-source line graph to explicitly model inter-document logical relationships, mitigating sparsity. Second, we design a dual-level confidence mechanism—operating at both graph and node levels—to dynamically identify and suppress inconsistent information. Our approach integrates graph-structured modeling, fine-grained knowledge fusion, and multi-level confidence-driven retrieval optimization. Evaluated on four multi-domain query datasets and two multi-hop question answering benchmarks, our method consistently outperforms state-of-the-art RAG approaches: hallucination rates decrease by 23.6% on average, while answer accuracy improves by 19.4%. The framework substantially enhances the reliability and robustness of multi-source knowledge retrieval in complex, real-world scenarios.
To address hallucinations in Retrieval-Augmented Generation (RAG) models caused by entanglement between parametric knowledge and externally retrieved knowledge, this paper conducts mechanistic interpretability analysis on the residual stream—revealing for the first time that hallucinations stem from excessive reliance of feed-forward networks (FFNs) on internal parametric knowledge and failure of copy heads to effectively integrate external content. Building on this insight, we propose a Knowledge Utilization Decoupling Detection paradigm and the Adaptive Activation Rescaling and Filtering (AARF) hallucination mitigation mechanism. Our approach models functional decoupling between FFNs and copy heads via targeted residual stream interventions and introduces an interpretable, lightweight hallucination detector. Evaluated across multiple benchmarks, our detector achieves an average 12.7% improvement in hallucination detection accuracy, while the AARF module reduces hallucination rates by up to 38.5%. Crucially, both components require no fine-tuning or additional training.
Existing adaptive retrieval-augmented generation (ARAG) systems excel at deep, single-source retrieval but struggle to anticipate and jointly regulate multi-source knowledge features, resulting in limited controllability and adaptability in cross-source retrieval. To address this, we propose MSPR—a Multi-Source Adaptive Retrieval-Augmented Generation framework—that pioneers joint decision-making on *when* to retrieve, *what* to retrieve, and *which source* to use. Its core innovations are: (1) a dual-track retrieval mechanism driven by chain-of-thought reasoning and retrieval preference modeling; (2) a dynamic retrieval-action optimization strategy guided by answer-level feedback; and (3) a complementary primary-secondary source selection paradigm. MSPR integrates reasoning-aware retrieval, preference-aware source modeling, and feedback-driven refinement. Evaluated on three benchmark datasets, MSPR achieves significant improvements in answer accuracy and factual consistency while substantially reducing hallucination rates, outperforming state-of-the-art ARAG methods across all metrics.
Retrieval-augmented generation (RAG) systems mitigate large language model hallucinations but remain vulnerable to adversarial corpus poisoning attacks, which induce factual errors in generated outputs. This paper presents the first systematic analysis of RAG’s two-stage failure mechanism, identifying retrieval ranking bias as the primary driver of successful attacks. To address this, we propose “skeptical prompting”—a lightweight, model-agnostic self-validation framework that operates at the generation stage without fine-tuning. It integrates multi-round consistency verification with knowledge activation assessment to enhance output robustness. Through retrieval quality attribution analysis and targeted adversarial sample construction, we conduct empirical validation across diverse benchmarks. Experimental results demonstrate that our approach reduces erroneous response rates by up to 47%, offering a practical, deployable defense for secure RAG systems.
This work addresses the persistent hallucination problem in retrieval-augmented generation (RAG) systems, which often occurs even when relevant documents are available, and highlights the limitations of conventional evaluation methods in identifying fine-grained issues in evidence utilization. The authors propose a diagnostic framework for question-answering tasks that decomposes questions into atomic reasoning facets and constructs a facet–text chunk matrix. By integrating retrieval relevance with natural language inference–based faithfulness scores, the framework analyzes evidence usage across three reasoning paradigms: Strict RAG, Soft RAG, and LLM-only. This approach enables, for the first time, facet-level diagnosis of RAG behavior, uncovering systematic failure modes—such as missing, misaligned, or prior-dominated evidence—that remain invisible at the answer level. Experimental results demonstrate that hallucinations primarily stem from flawed evidence integration rather than retrieval inaccuracies, offering an interpretable foundation for improving RAG systems.
This work addresses the risk of representational harm in retrieval-augmented generation (RAG) systems operating in high-stakes scenarios, where biases in the retrieval stage can propagate downstream. To mitigate this, the authors propose two exposure-aware ranking strategies: Forced-Exposure and Representative Stochastic. The latter explicitly acknowledges that initial relevance scores are already biased and instead seeks to achieve approximately fair group exposure, thereby moving beyond the limitations of conventional unbiasedness assumptions. Evaluated on the TREC 2022 Fair Ranking dataset with Wikipedia articles annotated into protected and non-protected categories, Representative Stochastic significantly improves average exposure fairness. Moreover, the demographic fairness of the generated outputs closely aligns with retrieval-stage exposure, underscoring the critical role of retrieval in controlling downstream bias.
This work addresses the compounding hallucination problem in retrieval-augmented generation (RAG) caused by erroneous retrieval. To mitigate this issue, the authors propose VOTE-RAG, a training-free, two-stage parallel voting framework. In the first stage, multiple agents independently generate queries and retrieve documents, which are then aggregated; in the second stage, answers are generated independently from the aggregated documents, with the final output determined by majority voting. This approach introduces a novel dual-voting mechanism that effectively suppresses hallucinations end-to-end without increasing model complexity or inducing query drift. Experimental results demonstrate that VOTE-RAG matches or outperforms more sophisticated existing methods across six benchmark datasets, highlighting the efficacy and superiority of this lightweight ensemble strategy in enhancing RAG reliability.
Existing multimodal RAG approaches for Visual Relation Detection (VRD) over-rely on salient textual and visual elements, neglecting fine-grained knowledge—such as small-font text and contextual cues—leading to incomplete retrieval and inaccurate answer generation. To address this, we propose SFT-RAG, the first framework to explicitly model both salient and fine-grained textual knowledge via a hybrid masking strategy. Additionally, we design an uncertainty-guided surrogate generator that dynamically fuses dual-path information flows. By integrating multimodal representation learning with dynamic knowledge integration, SFT-RAG achieves state-of-the-art performance on open-domain visual question answering: it attains top results under both zero-shot and supervised settings. Extensive experiments validate that explicit fine-grained knowledge modeling is critical for improving answer completeness and reliability.