Score
Designs and implements retrieval-augmented generation systems that restrict retrieval to a predefined set of candidate documents or passages (candidate-constrained or constrained RAG). Builds mechanisms that enforce generation to cite only those candidate documents, maintain deterministic provenance linking outputs to retrieved passages, and optimize answer quality under the constraint that only the provided candidates may be used.
This work addresses the challenge in LongEval-RAG tasks where responses must be strictly grounded in a given set of candidate documents. To this end, the authors propose a candidate-constrained retrieval-augmented generation (RAG) system that integrates rule-based chunking, query expansion, pseudo-relevance feedback, reciprocal rank fusion, MiniLM sentence-level reranking, and citation-aware evidence aggregation, complemented by deterministic provenance tracing and a neural sentence selection mechanism. Experimental results demonstrate that the proposed rule-MiniLM variant significantly outperforms baselines across multiple metrics—including BERTScore, retrieval precision, information point coverage, and human evaluation—thereby validating the effectiveness of combining rule-based chunking with neural sentence selection. The study further underscores the critical role of multi-metric evaluation in diagnosing and advancing RAG system performance.
This study systematically investigates the impact mechanisms of individual components in Retrieval-Augmented Generation (RAG) systems on complex question answering and cross-domain tasks. Addressing key challenges—including low retrieval precision, weak contextual relevance, and poor multilingual adaptability—we propose three core innovations: (1) a Contrastive In-Context Learning (CICL) RAG paradigm to improve generation accuracy; (2) sentence-granularity focused retrieval (“Focus Mode”) combined with multi-granularity chunking to enhance retrieval relevance; and (3) a multilingual knowledge base integration framework that balances retrieval–generation efficiency. Through large-scale hyperparameter analysis, we quantitatively characterize the influence of critical factors—including language model scale, chunk size, and retrieval stride—on end-to-end performance. The findings yield a reproducible best-practice guideline for RAG system design and deployment, accompanied by open-sourced, fully implemented code.
The impact of key design decisions—RAG activation, retrieval granularity, and knowledge integration strategy—on RAG system performance remains poorly understood. Method: We conduct systematic ablation studies across three code/qa benchmarks and two state-of-the-art LLMs, quantitatively evaluating how document type, recall rate, document selection strategy, and prompt engineering jointly affect answer correctness and confidence via multi-dimensional analysis, cross-model/dataset comparison, and joint prompt-retrieval analysis. Contribution/Results: We identify precise interaction patterns and operational boundaries among these factors and propose nine actionable, empirically grounded guidelines for diagnosing and optimizing RAG failures. Our findings significantly improve RAG system stability, debuggability, and reliability, offering rigorous empirical evidence and a principled methodology to support the engineering deployment of LLM-augmented systems.
This work addresses the challenges of ambiguous citation provenance and content redundancy commonly encountered in existing retrieval-augmented generation (RAG) systems during information integration. The authors propose a knowledge base construction approach grounded in Q&A nuggets, which leverages explicit question-answer semantics to guide information extraction, selection, and generation while preserving source attribution throughout the pipeline. Departing from conventional fuzzy clustering abstractions, the method employs interpretable Q&A fragments as structured intermediate representations, enabling end-to-end traceable reasoning and generation. Experimental results on the TREC NeuCLIR 2024 dataset demonstrate that the proposed approach significantly outperforms the state-of-the-art nugget-based RAG system, Ginger, in terms of nugget recall, density, and citation accuracy.
RAG system performance critically depends on the retriever-reader configuration, retrieval depth, and context quality; suboptimal settings often degrade performance rather than improve it. Method: We propose RAGGED, a systematic evaluation framework that—through multidimensional controlled experiments, controlled noise injection, and response attribution analysis—quantifies language models’ sensitivity spectra to contextual signals versus noise, and establishes a behavior-driven diagnostic paradigm for RAG configuration. Contribution/Results: We identify two canonical performance patterns—monotonic improvement and inverted-U—as context quality varies. Crucially, we reveal fundamental disparities across models in noise robustness and signal utilization capacity. Based on these insights, we distill reusable, model-aware configuration principles, validated across multiple DBQA benchmarks for both effectiveness and generalizability.
Traditional RAG systems suffer from suboptimal performance due to tight coupling among retrieval, reranking, prompt rewriting, and generation modules, hindering holistic optimization. Method: This paper proposes the first end-to-end RAG architecture search framework, modeling the RAG configuration space as an evolvable search problem and employing genetic algorithms to jointly optimize multi-objective metrics—including recall@k/nDCG for retrieval and LLM-Judge/semantic similarity for generation. The framework encompasses nine component types across vector retrieval, reranking, and prompt rewriting. Contribution/Results: Evaluated across six domains, the framework achieves an average 3.8% performance gain (up to +12.5% in retrieval, +7.5% in generation) while converging after exploring only 0.2% of the configuration space. It identifies robust architectural patterns transferable across datasets and quantifies how domain characteristics and question types systematically influence optimal configurations.
This study addresses a critical oversight in current retrieval-augmented generation (RAG) systems: their reliance on human-oriented document representations, which neglect the distinct representational needs of large language models as content consumers. Under fixed retrieval results, the authors systematically evaluate the impact of 14 document representation strategies—including selection, summarization, and rewriting—on question-answering accuracy across four generative models. Introducing answer retention rate as a novel metric to assess whether transformed documents preserve the correct answer, controlled experiments reveal for the first time that answer retention is the primary driver of generation accuracy, challenging prior assumptions that attributed performance gains to specific representational mechanisms. Notably, when answer retention is high, variations in wording, structure, length, or query dependence exert minimal influence on accuracy, underscoring that preserving answer information outweighs representational form.
This study addresses the challenge of accurately predicting the performance gain of Retrieval-Augmented Generation (RAG) over non-RAG approaches in question-answering tasks. To overcome limitations of existing prediction methods, the authors propose a novel supervised post-generation predictor that explicitly models the semantic relationships among the input question, retrieved passages, and the generated answer. The approach integrates multi-dimensional signals from pre-retrieval, post-retrieval, and post-generation stages and is optimized through end-to-end training. Experimental results demonstrate that the proposed method significantly outperforms current prediction strategies across multiple benchmarks, offering a reliable basis for deciding whether to activate RAG in practical systems.
This work proposes a retrieval-augmented generation (RAG) evaluation framework tailored to deep informational needs, addressing the challenges of reasoning and factual consistency posed by multi-sentence narrative queries in complex real-world scenarios. The framework introduces narrative-style multi-sentence queries and a multi-level evaluation protocol, enabling comprehensive assessment of system performance in terms of completeness, attributability, and consistency over the MS MARCO V2.1 corpus. Aggregating over 150 system submissions, the study demonstrates the feasibility of simultaneously ensuring factual accuracy, transparency, and high-quality generation under complex querying conditions, thereby establishing a new benchmark for trustworthy, context-aware RAG systems.
This work addresses the challenge of leveraging LaTeX source code—rich in structural and semantic information yet hindered by cross-references, custom macros, and unmarked content—for retrieval-augmented generation (RAG). The authors propose a systematic preprocessing pipeline that transforms raw LaTeX documents and their auxiliary files into structured Markdown and JSONL fragments through LaTeX parsing, macro expansion, reference resolution, and semantic annotation. This pipeline enables efficient indexing in vector databases and constitutes the first end-to-end method to convert native LaTeX into a knowledge format readily consumable by large language models. By preserving document structure, semantic labels, and authorial intent, the approach significantly enhances the accuracy and reliability of language models on mathematical and technical question-answering tasks.