Score
Designs and applies methods, tests, and metrics to assess the accuracy, provenance, and trustworthiness of responses produced by retrieval-augmented generation (RAG) systems and graph-augmented LLM outputs. This work includes attributing claims to parametric versus retrieved sources, checking relevance and consistency of retrieved evidence, detecting inaccuracies or contradictions in knowledge graphs, and calibrating or flagging model confidence and trust for downstream use.
This study addresses the complexity of evaluating Retrieval-Augmented Generation (RAG) systems—stemming from their multi-component architecture (indexing, retrieval, generation) and high-dimensional parameter space, which hinder quantitative quality assessment. Through a systematic review of 63 papers, we construct a full-stack evaluation taxonomy spanning four dimensions: datasets, retrievers, indexes/databases, and generators. We pioneer the dual role of large language models (LLMs) in RAG evaluation: (i) automating the construction of high-quality, diverse test suites, and (ii) performing interpretable, multi-faceted assessments—including retrieval relevance, answer faithfulness, and information completeness. We rigorously delineate the boundaries and synergies between LLM-based automated evaluation and human judgment. Empirical validation across multiple benchmarks confirms the framework’s effectiveness, and we deliver a reusable, production-ready RAG evaluation guideline for industry practitioners.
Evaluating retrieval-augmented generation (RAG) systems in the large language model (LLM) era faces unique challenges—including hybrid architectures, dynamic knowledge sources, and multidimensional trustworthiness requirements. Method: We bridge traditional and LLM-native evaluation paradigms for the first time, proposing a four-dimensional taxonomy covering performance, factual accuracy, safety, and computational efficiency. Our meta-analysis synthesizes insights from 120+ high-impact studies. Evaluation methodology integrates human assessment, automated metrics (e.g., ROUGE, BERTScore), fact-checking tools, retrieval-quality measures, and end-to-end interpretability analysis. We further construct a RAG-specific benchmark dataset and a framework classification taxonomy to expose current practice biases. Contribution/Results: This work delivers the first standardized, multidimensional, trust-oriented evaluation guideline for RAG systems—enabling rigorous, reproducible, and responsible development and deployment of trustworthy RAG applications.
This paper addresses the lack of trustworthiness—specifically in robustness, privacy, adversarial resilience, and accountability—in Retrieval-Augmented Generation (RAG) systems. To this end, it introduces, for the first time, a unified six-dimensional analytical framework and taxonomy covering reliability, privacy, security, fairness, explainability, and accountability. The proposed multidisciplinary evaluation paradigm integrates principles from trustworthy AI, information retrieval, adversarial robustness, differential privacy, eXplainable AI (XAI), and responsibility modeling—thereby filling a critical gap in systematic RAG trustworthiness assessment. Furthermore, the work establishes an industrial-deployment-oriented RAG trustworthiness roadmap, explicitly identifying key technical bottlenecks and evolutionary pathways across all six dimensions. This provides both theoretical foundations and practical guidance for developing high-assurance AI-generated content (AIGC) systems.
This study investigates the factual reliability of Retrieval-Augmented Generation (RAG) systems when exposed to misleading information in retrieved passages. By constructing controlled test scenarios comprising clean, contaminated, and mixed evidence, and integrating factual question-answering benchmarks with comparative analyses between parametric knowledge and retrieved evidence, the work proposes an evaluation framework that combines parametric coverage and confidence metrics. It systematically uncovers, for the first time, the mechanisms through which misinformation influences large language model generation, quantifies RAG’s vulnerability under conflicting information, and provides a reproducible methodology along with empirical evidence to enhance its robustness.
This study addresses the declining factual accuracy of Retrieval-Augmented Generation (RAG) in high-stakes medical applications, where outdated or contradictory source documents undermine reliability. We introduce a novel medical question-answering benchmark built upon Australian Therapeutic Goods Administration (TGA) product labeling and propose time-stratified PubMed abstract retrieval to enable controlled evaluation of outdated evidence. Experiments compare five state-of-the-art LLMs on integrating temporally dispersed, semantically similar yet contradictory medical literature. Results demonstrate that lexical or semantic similarity alone does not ensure RAG reliability in medicine; contradictory evidence significantly degrades both answer accuracy and inter-response consistency. Our key contribution is the empirical revelation that “similarity ≠ reliability” — a critical limitation in medical RAG — and the first systematic validation of the necessity of contradiction-aware filtering mechanisms. This work provides empirical grounding and methodological guidance for enhancing RAG safety in high-risk domains.
To address the low trustworthiness of large language models (LLMs) in retrieval-augmented generation (RAG) systems, this paper proposes Trust-Score—a quantitative framework for multi-dimensional trust assessment—and Trust-Align—a lightweight, fine-tuning-free alignment method. Trust-Score jointly models factual citation accuracy, refusal capability, and attribution grounding to holistically quantify LLM trustworthiness. Trust-Align enhances robustness in unanswerable question detection and evidence attribution through learned refusal mechanisms, multi-task prompt alignment, and cross-model generalization adaptation. Evaluated on ASQA, QAMPARI, and ELI5 benchmarks, our approach outperforms 26 out of 27 open-source models—including a +12.56% improvement for LLaMA-3-8B on ASQA—and efficiently adapts to models ranging from 1B to 8B parameters. To our knowledge, this is the first work to establish an interpretable, generalizable, and fine-tuning-free paradigm for enhancing LLM trustworthiness in RAG settings.
Retrieval-Augmented Generation (RAG) systems exhibit insufficient reliability in high-stakes domains—such as investment due diligence—where hallucinations, off-topic responses, and citation failures pose critical risks. Method: We propose a statistically rigorous yet industrially scalable evaluation framework that integrates human expert annotations with calibrated LLM-based judges in a collaborative assessment protocol. Inspired by predictive inference, our approach enables precise, quantitative identification of multiple failure modes—including hallucination, irrelevance, and retrieval failure. Contributions/Results: (1) We introduce the first RAG reliability benchmark specifically designed for fund due diligence; (2) we publicly release a high-quality evaluation dataset and open-source evaluation code; (3) empirical validation demonstrates that our framework significantly improves assessment reliability while enabling large-scale automated evaluation—establishing a reproducible, verifiable evaluation paradigm for deploying RAG systems in mission-critical applications.
To address low credibility, poor traceability, and insufficient answer diversity of Retrieval-Augmented Generation (RAG) in low-resource domain expert systems handling heterogeneous multimodal data, this paper proposes an end-to-end trustworthy RAG framework. First, it introduces a structured corpus construction and Q&A auto-generation pipeline tailored for messy, real-world data. Second, it designs a semantic- and evidence-confidence-driven two-stage re-ranking mechanism to enhance retrieval precision. Third, it pioneers an LLM-guided reference matching algorithm that ensures answer grounding in retrieved evidence and enables fully traceable generation. Experiments in automotive engineering demonstrate significant improvements over non-RAG baselines: +1.94 in factual correctness, +1.16 in informativeness, and +1.67 in helpfulness (5-point scale, evaluated by LLM judges), validating the framework’s effectiveness in trustworthy retrieval, interpretable generation, and auditability.
This work addresses the issue of factual inconsistency between generated answers and cited sources in retrieval-augmented generation (RAG) systems by proposing a corrective RAG pipeline that integrates pre-generation passage filtering with post-generation strict entailment verification. Building upon Corrective RAG and CiteFix mechanisms, the approach further incorporates an LLM-as-judge diagnostic method to enhance citation fidelity and factual grounding. The proposed framework effectively improves the faithfulness of citations while preserving answer relevance and fluency, thereby demonstrating the feasibility of strengthening citation integrity in RAG outputs. Moreover, the study advocates for a new evaluation paradigm that prioritizes strict answer traceability to source evidence, emphasizing the need for more rigorous assessment of attribution accuracy in generative retrieval systems.
This work addresses a critical limitation in current retrieval-augmented generation (RAG) fact-checking systems, which typically assume retrieved evidence is reliable while overlooking that sources may be biased, unreliable, or outdated. To mitigate this, the authors introduce MEDIAREF—the first publicly available and updatable media background knowledge base, encompassing 200 news outlets. Constructed through web crawling and document aggregation, MEDIAREF enables Media Background Checking (MBC) via large language models without reliance on costly proprietary APIs. Experimental results demonstrate that integrating MEDIAREF significantly enhances models’ ability to assess source credibility, thereby improving the transparency and reliability of fact-checking systems. Furthermore, MEDIAREF advances reproducible research and equitable evaluation in source-critical reasoning.
Incorporating specific knowledge into large language models via retrieval-augmented generation (RAG) is a widespread technique that fuels many of today's industry AI applications. A fundamental problem is to assess if the context retrieved by some similarity search provides indeed supporting facts, or instead misguides the generator with irrelevant information. It is critical to associate meaningful confidence measures about the factuality of the retrieval process with the generated answers. We present a new, two-staged approach to predict fact faithfulness of the output of retrieval-augmented generations. First, we employ conformal prediction to select only those retrieved chunks who have a high chance to come from the correct source. This approach in itself can improve answer quality by up to 6% in some of the studied datasets, however, the associated statistical guarantees do not hold generally, since the assumption of sample exchangeability depends on the retriever setup. We present diagnostic metrics to assess whether a setup is suitable. Second, we quantify confidence in the consistency of a generated final answer with a given retrieved context, using an attention-based factuality classifier. This approach can detect inconsistent answers with a chance of up to 77%. Our work helps to establish a novel type of certified RAG systems for a broad range of natural language industry applications.