Score
Designs and implements retrieval pipelines that search a vetted, curated corpus to extract evidence-backed answers and attach explicit source citations; includes query formulation, citation selection, and presentation logic to surface the most relevant vetted documents. Builds and analyzes trust and relevance filters and refusal criteria so the system reduces untrustworthy or irrelevant citations and appropriately declines to answer when the curated corpus provides insufficient support.
This study addresses the inherent trade-off between coverage breadth and information reliability in public AI information services by empirically comparing retrieval-augmented generation (RAG) systems built on curated corpora against open web search in the context of EU governmental question-answering. Through expert evaluation, it identifies “source credibility” as a latent yet quantifiable dimension of response quality. Findings reveal that while open web search achieves broader coverage, 35% of its cited sources are either untrustworthy or irrelevant; conversely, curated corpora ensure high source reliability but suffer from limited scope. Moreover, system prompts exhibit minimal efficacy in steering models toward citing trustworthy domains. These results provide empirical grounding and practical guidance for designing information-source strategies in public-facing AI systems.
This work proposes an end-to-end framework to assist users in evaluating the credibility of online news. The system first leverages a large language model (LLM) to generate diverse, critical questions about a given news claim. It then retrieves relevant evidence from large-scale corpora such as MS MARCO V2.1 through semantic filtering, clustering, and a novel Chain-of-Thought query expansion strategy. Subsequently, retrieved passages are re-ranked using monoT5 and further assessed for relevance via LLM-driven judgment, culminating in a citation-backed credibility report. Experimental results demonstrate that the proposed Chain-of-Thought query expansion and re-ranking mechanisms significantly enhance both evidence relevance and domain-specific credibility, validating the efficacy of the approach, while also indicating room for improvement in the quality of generated questions.
This work addresses the challenge in LongEval-RAG tasks where responses must be strictly grounded in a given set of candidate documents. To this end, the authors propose a candidate-constrained retrieval-augmented generation (RAG) system that integrates rule-based chunking, query expansion, pseudo-relevance feedback, reciprocal rank fusion, MiniLM sentence-level reranking, and citation-aware evidence aggregation, complemented by deterministic provenance tracing and a neural sentence selection mechanism. Experimental results demonstrate that the proposed rule-MiniLM variant significantly outperforms baselines across multiple metrics—including BERTScore, retrieval precision, information point coverage, and human evaluation—thereby validating the effectiveness of combining rule-based chunking with neural sentence selection. The study further underscores the critical role of multi-metric evaluation in diagnosing and advancing RAG system performance.
To address insufficient evidence coverage and low answer accuracy in Retrieval-Augmented Generation (RAG) for government document question answering within legal and regulatory domains, this paper proposes two synergistic optimization strategies. First, a One-SHOT retrieval method with adaptive token budgeting improves recall of critical information chunks. Second, an iterative retrieval framework built upon a Reasoning Agentic RAG architecture integrates dynamic query generation, progressive context refinement, and result evaluation with feedback—effectively mitigating query drift and retrieval inertia. Experimental results demonstrate substantial improvements: +28.6% in evidence coverage and +14.3 BLEU points in answer accuracy. These advances establish a novel paradigm for high-precision, interpretable legal intelligent question answering.