Score
Designs and implements retrieval-augmented generation pipelines that integrate compact/small language models with retriever outputs, including mechanisms to incorporate retrieved context into prompts or encoders and to produce grounded answers. Builds evaluations and benchmarks of generation quality and latency–quality tradeoffs across datasets, analyzing how retrieval, context selection, and model size/compression affect accuracy, fluency, and system throughput.
This study investigates the capacity of small language models (7B parameters or fewer) to effectively leverage external information in retrieval-augmented generation (RAG). Through systematic evaluation on models such as SmolLM2, Qwen2.5, and Llama 3.1—combined with BM25, E5-large-v2, and oracle retrievers across multiple prompt templates—the work introduces a novel parameterized knowledge partitioning framework that cleanly disentangles retrieval failure from context utilization failure for the first time. The findings reveal a fundamental bottleneck in small models’ ability to use retrieved content: even under oracle retrieval conditions, 85%–100% of samples fail to correctly incorporate the relevant answer, and 42%–100% of the model’s original knowledge is disrupted by the retrieved context. The dominant error mode is generation entirely unrelated to the provided context, indicating a pervasive inability to attend to or integrate external information.
This study systematically investigates the impact mechanisms of individual components in Retrieval-Augmented Generation (RAG) systems on complex question answering and cross-domain tasks. Addressing key challenges—including low retrieval precision, weak contextual relevance, and poor multilingual adaptability—we propose three core innovations: (1) a Contrastive In-Context Learning (CICL) RAG paradigm to improve generation accuracy; (2) sentence-granularity focused retrieval (“Focus Mode”) combined with multi-granularity chunking to enhance retrieval relevance; and (3) a multilingual knowledge base integration framework that balances retrieval–generation efficiency. Through large-scale hyperparameter analysis, we quantitatively characterize the influence of critical factors—including language model scale, chunk size, and retrieval stride—on end-to-end performance. The findings yield a reproducible best-practice guideline for RAG system design and deployment, accompanied by open-sourced, fully implemented code.
This work addresses key challenges in deploying large language models for retrieval-augmented generation (RAG), including high computational overhead, rapid knowledge obsolescence, and manual dependency in component selection. The authors propose a modular evaluation framework that, for the first time, directly links hardware constraints to RAG performance. By integrating resource telemetry with an automated recommendation mechanism, the framework efficiently identifies optimal combinations of components—including document chunking strategies, embedding models, vector databases, and retrievers—for domain-specific datasets. This approach maintains high generation quality while substantially reducing resource consumption. Designed to support rapid prototyping on consumer-grade hardware, the framework enables automatic, domain-tailored RAG configuration, achieving a favorable trade-off among accuracy, efficiency, and scalability.
This study addresses the lack of systematic evaluation of small language models (SLMs) in retrieval-augmented generation (RAG) systems, particularly their potential for deployment on resource-constrained devices. It presents the first comprehensive assessment of SLMs in the RAG generation phase, benchmarking performance across diverse domains using both open-source and proprietary datasets. The work further demonstrates end-side inference entirely on CPU-based hardware without GPU acceleration. Experimental results show that SLMs can operate efficiently under such constraints, substantially reducing computational overhead while maintaining reasonable response times and generating high-quality outputs. These findings establish a viable pathway for deploying lightweight, edge-compatible AI systems leveraging RAG architectures.
The impact of key design decisions—RAG activation, retrieval granularity, and knowledge integration strategy—on RAG system performance remains poorly understood. Method: We conduct systematic ablation studies across three code/qa benchmarks and two state-of-the-art LLMs, quantitatively evaluating how document type, recall rate, document selection strategy, and prompt engineering jointly affect answer correctness and confidence via multi-dimensional analysis, cross-model/dataset comparison, and joint prompt-retrieval analysis. Contribution/Results: We identify precise interaction patterns and operational boundaries among these factors and propose nine actionable, empirically grounded guidelines for diagnosing and optimizing RAG failures. Our findings significantly improve RAG system stability, debuggability, and reliability, offering rigorous empirical evidence and a principled methodology to support the engineering deployment of LLM-augmented systems.
This study addresses a critical oversight in current retrieval-augmented generation (RAG) systems: their reliance on human-oriented document representations, which neglect the distinct representational needs of large language models as content consumers. Under fixed retrieval results, the authors systematically evaluate the impact of 14 document representation strategies—including selection, summarization, and rewriting—on question-answering accuracy across four generative models. Introducing answer retention rate as a novel metric to assess whether transformed documents preserve the correct answer, controlled experiments reveal for the first time that answer retention is the primary driver of generation accuracy, challenging prior assumptions that attributed performance gains to specific representational mechanisms. Notably, when answer retention is high, variations in wording, structure, length, or query dependence exert minimal influence on accuracy, underscoring that preserving answer information outweighs representational form.
This study addresses the lack of systematic, up-to-date synthesis of retrieval-augmented generation (RAG) research amid rapid methodological diversification and evaluation fragmentation. Adopting the PRISMA 2020 framework, we systematically curated 128 highly cited papers (2020–2025) from ACM, IEEE, and other authoritative sources, introducing a dynamic citation threshold to mitigate temporal bias. Our analysis maps evolutionary trajectories across three dimensions: architectural design, benchmark datasets, and evaluation metrics—revealing critical methodological gaps, particularly in non-parametric memory augmentation and neural retrieval–generation co-adaptation. We construct a structured knowledge graph of RAG research and propose a prioritized roadmap that jointly optimizes robustness, interpretability, and generalization. The findings provide empirically grounded guidance for both foundational RAG theory development and practical deployment.
This work addresses the lack of systematic understanding of the differences and complementarities among diverse retrievers in current Retrieval-Augmented Generation (RAG) systems, which hinders effective retriever selection and integration. To this end, we propose MIGRASCOPE, a novel framework that introduces mutual information and statistical estimation theory into RAG evaluation, enabling quantitative assessment of retrievers in terms of retrieval quality, redundancy, synergy, and marginal contribution. Leveraging this framework, we uncover complementary relationships among mainstream retrievers and design efficient integration strategies. Experimental results demonstrate that carefully orchestrated multi-retriever systems significantly outperform the best single retriever, offering both theoretical grounding and practical guidance for building more effective and robust RAG systems.
Existing asynchronous retrieval-augmented generation (RAG) systems rely on heuristic coordination strategies that struggle to adapt to dynamically evolving information needs across diverse domains, resulting in limited efficiency and flexibility. To address this, this work proposes a novel asynchronous retrieval framework that leverages semantic precursors emerging early in the generation process to explicitly predict both the optimal timing and content for retrieval, enabling intelligent prefetching aligned with dynamic user demands. The framework integrates a retrieval predictor, a context monitor, and a query generator, jointly modeling semantic precursors and evolving information requirements. Experimental results demonstrate that the approach reduces end-to-end latency by up to 43.5% and accelerates first-token output speed by up to 62.4% across multiple benchmarks, while maintaining answer quality comparable to that of synchronous RAG systems.