Score
Designs and evaluates retrieval-augmented generation systems that use observed errors or selected strategies to retrieve, rank, and integrate repair examples or strategy snippets into generation and decision logic. This competence covers mapping diagnostic signals to candidate fixes, building retrieval and prioritization modules for repair content or strategies, and conditioning the generator or repair policy on retrieved items to produce or recommend corrections.
This paper addresses two critical challenges in retrieval-augmented generation (RAG): the lack of clarity regarding the synergistic mechanisms between RAG and large language model (LLM) reasoning capabilities, and the absence of a comprehensive evaluation framework. To this end, we propose the first unified taxonomy for RAG-Reasoning synergy, systematically defining “reasoning under RAG” across three dimensions—collaborative objectives, canonical paradigms, and technical implementations—and analyzing bidirectional synergy pathways. Through a critical evaluation, we identify key blind spots in current RAG benchmarks, notably the absence of intermediate reasoning supervision and insufficient cost-effectiveness trade-off analysis. We further introduce three novel research directions: knowledge graph integration, hybrid-model collaborative reasoning, and reinforcement learning–driven optimization. Our work establishes the first theoretically grounded and practically actionable RAG-Reasoning synergy framework, providing foundational support for academic standardization and industrial-scale RAG system advancement.
To address the limitations of existing RAG systems in complex reasoning, dynamic retrieval, and multimodal integration within real-world industrial applications, this paper proposes an inference-enhanced intelligent RAG framework. Methodologically, it introduces the first dual-track reasoning taxonomy—System 1 (fast, modular reasoning) and System 2 (slow, autonomous planning)—and establishes the first open-source knowledge-graph-based RAG survey repository. The framework integrates LLM-driven reasoning architectures, standardized tool-use protocols (e.g., ReAct), multi-stage retrieval strategies, and multimodal interfaces. Through a systematic analysis of over 120 state-of-the-art works, we identify seven inference patterns and five solutions to key industrial bottlenecks. Empirical evaluation in production scenarios—including customer service and financial risk control—demonstrates 23%–38% improvements in reasoning accuracy.
To address the low efficiency and poor accuracy of information retrieval and maintenance instruction generation for technicians handling heterogeneous multimodal data (e.g., text, images, 3D models) in XR environments, this paper proposes the first cross-format Retrieval-Augmented Generation (RAG) framework tailored for industrial XR. The framework achieves unified retrieval via cross-modal semantic alignment and integrates large language models (LLMs)—specifically GPT-4 and GPT-4o-mini—to generate context-aware maintenance instructions. Its key innovation lies in the first end-to-end integration of joint multimodal retrieval and LLM-based instruction generation within an XR runtime. Experimental results demonstrate a 37% improvement in instruction response accuracy and an average latency of 1.18 seconds; for complex queries, BLEU and METEOR scores reach 42.6 and 48.3, respectively—validating the framework’s superior real-time performance, accuracy, and industrial applicability.
To address error propagation and answer bias arising from unidirectional retrieval-then-reasoning in multi-hop question answering, this paper proposes RetroRAG, the first framework introducing backtracking-style reasoning. Its core is an evidence backtracking mechanism: inferring entity-centric queries to dynamically revise retrieved evidence and reconstruct reasoning paths, enabling iterative refinement and dynamic reorganization of trustworthy evidence through coordinated multi-round retrieval-generation-evaluation cycles. This establishes a closed-loop “evidence curation–discovery–verification” process, substantially enhancing robustness and interpretability for complex reasoning. On mainstream multi-hop QA benchmarks, RetroRAG consistently outperforms existing RAG methods, achieving significant gains in answer accuracy—particularly under challenging conditions involving long reasoning chains and noisy evidence.
This study addresses the challenge faced by operators in industrial settings who struggle to rapidly locate relevant troubleshooting procedures from vast volumes of technical documentation matching specific fault symptoms. To tackle this issue, the work proposes a retrieval-augmented generation (RAG)-based conversational assistance system, which is validated for the first time in a large-scale maritime cyber-physical system to demonstrate RAG’s practical efficacy in complex fault diagnosis scenarios. Experimental results show that the proposed approach significantly improves both the speed and accuracy of operator responses. Furthermore, the study underscores the necessity of incorporating cross-validation mechanisms to ensure the reliability of AI-generated recommendations, thereby offering actionable guidelines for deploying trustworthy AI-assisted decision-making in high-risk industrial environments.
The impact of key design decisions—RAG activation, retrieval granularity, and knowledge integration strategy—on RAG system performance remains poorly understood. Method: We conduct systematic ablation studies across three code/qa benchmarks and two state-of-the-art LLMs, quantitatively evaluating how document type, recall rate, document selection strategy, and prompt engineering jointly affect answer correctness and confidence via multi-dimensional analysis, cross-model/dataset comparison, and joint prompt-retrieval analysis. Contribution/Results: We identify precise interaction patterns and operational boundaries among these factors and propose nine actionable, empirically grounded guidelines for diagnosing and optimizing RAG failures. Our findings significantly improve RAG system stability, debuggability, and reliability, offering rigorous empirical evidence and a principled methodology to support the engineering deployment of LLM-augmented systems.
This work addresses the challenges of ambiguous citation provenance and content redundancy commonly encountered in existing retrieval-augmented generation (RAG) systems during information integration. The authors propose a knowledge base construction approach grounded in Q&A nuggets, which leverages explicit question-answer semantics to guide information extraction, selection, and generation while preserving source attribution throughout the pipeline. Departing from conventional fuzzy clustering abstractions, the method employs interpretable Q&A fragments as structured intermediate representations, enabling end-to-end traceable reasoning and generation. Experimental results on the TREC NeuCLIR 2024 dataset demonstrate that the proposed approach significantly outperforms the state-of-the-art nugget-based RAG system, Ginger, in terms of nugget recall, density, and citation accuracy.
This study addresses the challenges of decision latency and information overload in space operations caused by the vast volume of technical documentation and scientific literature. It presents the first systematic evaluation of Retrieval-Augmented Generation (RAG) for this high-stakes domain, integrating multiple retrieval strategies, embedding models, and large language models to efficiently extract and synthesize actionable knowledge from domain-specific documents. Experimental results demonstrate that the proposed RAG pipeline substantially enhances the accuracy, relevance, and reliability of knowledge retrieval, thereby reducing decision uncertainty. The work delivers a practical and trustworthy intelligent support framework for complex space missions while delineating clear pathways for optimization and defining the boundaries of its applicability.
Current RAG system evaluations overly rely on end-to-end accuracy, failing to capture enterprise-level requirements across dimensions such as reasoning complexity, retrieval difficulty, document structural diversity, and interpretability. Consequently, models achieving high scores often exhibit insufficient reliability in real-world deployments. To address this gap, this work proposes the first difficulty taxonomy integrating these four dimensions and introduces a multidimensional diagnostic framework and benchmark tailored for enterprise applications. The framework systematically identifies weaknesses of RAG systems in complex, realistic settings and effectively exposes performance bottlenecks that hinder practical deployment, thereby offering actionable pathways for evaluation and optimization to enhance real-world reliability.
This work addresses the lack of evaluation frameworks for assessing how retrieval-augmented generation (RAG) systems adapt following user or expert feedback. It introduces, for the first time, a “feedback adaptation” problem setting, quantifying adaptation speed and reliability through two metrics: correction latency and post-feedback performance. To enable real-time feedback integration without retraining during inference, the authors propose PatchRAG, which combines semantic relevance analysis with behavioral change detection to achieve zero-latency corrections and cross-query semantic generalization. Experimental results demonstrate that PatchRAG significantly outperforms baseline methods, maintaining immediate responsiveness while exhibiting strong generalization capabilities after receiving feedback.
This work addresses the prevalence of factual errors in deployed Retrieval-Augmented Generation (RAG) systems, which often stem from missing retrieval evidence or contextually inconsistent generation—issues that existing repair methods struggle to resolve under black-box conditions or resource constraints. To tackle this challenge, the authors propose D2R-RAG, a novel model-agnostic and resource-aware framework for diagnosing and repairing RAG failures. D2R-RAG employs lightweight modules to extract interpretable failure signatures from the query, retrieved passages, and generated response, then adaptively selects the optimal repair strategy under explicit latency and memory budgets. Experimental results demonstrate that D2R-RAG significantly enhances reliability on FEVER and HotpotQA, consistently outperforming existing baselines across diverse computational budgets while achieving a superior trade-off between accuracy and efficiency.