error-guided rag

Designs and evaluates retrieval-augmented generation systems that use observed errors or selected strategies to retrieve, rank, and integrate repair examples or strategy snippets into generation and decision logic. This competence covers mapping diagnostic signals to candidate fixes, building retrieval and prioritization modules for repair content or strategies, and conditioning the generator or repair policy on retrieved items to produce or recommend corrections.

error-guidedrag

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Reasoning RAG via System 1 or System 2: A Survey on Reasoning Agentic Retrieval-Augmented Generation for Industry Challenges

Jun 12, 2025
JL
Jintao Liang
🏛️ Beijing University of Posts and Telecommunications | University of Georgia | South China University of Technology | SenseTime Research | Qingyuan Research Institute | Shanghai Jiaotong University | Technical University of Munich | University of Cologne

To address the limitations of existing RAG systems in complex reasoning, dynamic retrieval, and multimodal integration within real-world industrial applications, this paper proposes an inference-enhanced intelligent RAG framework. Methodologically, it introduces the first dual-track reasoning taxonomy—System 1 (fast, modular reasoning) and System 2 (slow, autonomous planning)—and establishes the first open-source knowledge-graph-based RAG survey repository. The framework integrates LLM-driven reasoning architectures, standardized tool-use protocols (e.g., ReAct), multi-stage retrieval strategies, and multimodal interfaces. Through a systematic analysis of over 120 state-of-the-art works, we identify seven inference patterns and five solutions to key industrial bottlenecks. Empirical evaluation in production scenarios—including customer service and financial risk control—demonstrates 23%–38% improvements in reasoning accuracy.

Advancing agentic RAG with autonomous tool interactionEnhancing RAG for complex reasoning and dynamic retrievalOvercoming knowledge limitations in LLMs via RAG

Must-Read Papers

Most classic and influential ideas
View more

To address the low efficiency and poor accuracy of information retrieval and maintenance instruction generation for technicians handling heterogeneous multimodal data (e.g., text, images, 3D models) in XR environments, this paper proposes the first cross-format Retrieval-Augmented Generation (RAG) framework tailored for industrial XR. The framework achieves unified retrieval via cross-modal semantic alignment and integrates large language models (LLMs)—specifically GPT-4 and GPT-4o-mini—to generate context-aware maintenance instructions. Its key innovation lies in the first end-to-end integration of joint multimodal retrieval and LLM-based instruction generation within an XR runtime. Experimental results demonstrate a 37% improvement in instruction response accuracy and an average latency of 1.18 seconds; for complex queries, BLEU and METEOR scores reach 42.6 and 48.3, respectively—validating the framework’s superior real-time performance, accuracy, and industrial applicability.

Enhance information retrieval for maintenanceIntegrate multi-format data with LLMsOptimize response accuracy and speed

Retrieval-Augmented Generation by Evidence Retroactivity in LLMs

Jan 07, 2025
LX
Liang Xiao
🏛️ Beijing Institute of Technology | Xiaomi Corporation

To address error propagation and answer bias arising from unidirectional retrieval-then-reasoning in multi-hop question answering, this paper proposes RetroRAG, the first framework introducing backtracking-style reasoning. Its core is an evidence backtracking mechanism: inferring entity-centric queries to dynamically revise retrieved evidence and reconstruct reasoning paths, enabling iterative refinement and dynamic reorganization of trustworthy evidence through coordinated multi-round retrieval-generation-evaluation cycles. This establishes a closed-loop “evidence curation–discovery–verification” process, substantially enhancing robustness and interpretability for complex reasoning. On mainstream multi-hop QA benchmarks, RetroRAG consistently outperforms existing RAG methods, achieving significant gains in answer accuracy—particularly under challenging conditions involving long reasoning chains and noisy evidence.

Accuracy ImprovementInformation RetrievalLarge Language Models

This study addresses the challenge faced by operators in industrial settings who struggle to rapidly locate relevant troubleshooting procedures from vast volumes of technical documentation matching specific fault symptoms. To tackle this issue, the work proposes a retrieval-augmented generation (RAG)-based conversational assistance system, which is validated for the first time in a large-scale maritime cyber-physical system to demonstrate RAG’s practical efficacy in complex fault diagnosis scenarios. Experimental results show that the proposed approach significantly improves both the speed and accuracy of operator responses. Furthermore, the study underscores the necessity of incorporating cross-validation mechanisms to ensure the reliability of AI-generated recommendations, thereby offering actionable guidelines for deploying trustworthy AI-assisted decision-making in high-risk industrial environments.

cyber-physical systemfailure resolutioninformation retrieval

Towards Understanding Retrieval Accuracy and Prompt Quality in RAG Systems

Nov 29, 2024
SZ
Shengming Zhao
🏛️ University of Alberta | The University of Tokyo | East China Normal University

The impact of key design decisions—RAG activation, retrieval granularity, and knowledge integration strategy—on RAG system performance remains poorly understood. Method: We conduct systematic ablation studies across three code/qa benchmarks and two state-of-the-art LLMs, quantitatively evaluating how document type, recall rate, document selection strategy, and prompt engineering jointly affect answer correctness and confidence via multi-dimensional analysis, cross-model/dataset comparison, and joint prompt-retrieval analysis. Contribution/Results: We identify precise interaction patterns and operational boundaries among these factors and propose nine actionable, empirically grounded guidelines for diagnosing and optimizing RAG failures. Our findings significantly improve RAG system stability, debuggability, and reliability, offering rigorous empirical evidence and a principled methodology to support the engineering deployment of LLM-augmented systems.

Analyzing key engineering trade-offs in RAG deployment decisionsDetermining optimal retrieval volume for different task typesEvaluating effective knowledge integration methods across tasks

This work addresses the challenges of ambiguous citation provenance and content redundancy commonly encountered in existing retrieval-augmented generation (RAG) systems during information integration. The authors propose a knowledge base construction approach grounded in Q&A nuggets, which leverages explicit question-answer semantics to guide information extraction, selection, and generation while preserving source attribution throughout the pipeline. Departing from conventional fuzzy clustering abstractions, the method employs interpretable Q&A fragments as structured intermediate representations, enabling end-to-end traceable reasoning and generation. Experimental results on the TREC NeuCLIR 2024 dataset demonstrate that the proposed approach significantly outperforms the state-of-the-art nugget-based RAG system, Ginger, in terms of nugget recall, density, and citation accuracy.

citation provenanceinformation redundancyinterpretable generation

Latest Papers

What's happening recently
View more

This study addresses the challenges of decision latency and information overload in space operations caused by the vast volume of technical documentation and scientific literature. It presents the first systematic evaluation of Retrieval-Augmented Generation (RAG) for this high-stakes domain, integrating multiple retrieval strategies, embedding models, and large language models to efficiently extract and synthesize actionable knowledge from domain-specific documents. Experimental results demonstrate that the proposed RAG pipeline substantially enhances the accuracy, relevance, and reliability of knowledge retrieval, thereby reducing decision uncertainty. The work delivers a practical and trustworthy intelligent support framework for complex space missions while delineating clear pathways for optimization and defining the boundaries of its applicability.

decision-makinginformation retrievalknowledge access

Current RAG system evaluations overly rely on end-to-end accuracy, failing to capture enterprise-level requirements across dimensions such as reasoning complexity, retrieval difficulty, document structural diversity, and interpretability. Consequently, models achieving high scores often exhibit insufficient reliability in real-world deployments. To address this gap, this work proposes the first difficulty taxonomy integrating these four dimensions and introduces a multidimensional diagnostic framework and benchmark tailored for enterprise applications. The framework systematically identifies weaknesses of RAG systems in complex, realistic settings and effectively exposes performance bottlenecks that hinder practical deployment, thereby offering actionable pathways for evaluation and optimization to enhance real-world reliability.

enterprise benchmarkmulti-dimensional evaluationpractical deployment

This work addresses the lack of evaluation frameworks for assessing how retrieval-augmented generation (RAG) systems adapt following user or expert feedback. It introduces, for the first time, a “feedback adaptation” problem setting, quantifying adaptation speed and reliability through two metrics: correction latency and post-feedback performance. To enable real-time feedback integration without retraining during inference, the authors propose PatchRAG, which combines semantic relevance analysis with behavioral change detection to achieve zero-latency corrections and cross-query semantic generalization. Experimental results demonstrate that PatchRAG significantly outperforms baseline methods, maintaining immediate responsiveness while exhibiting strong generalization capabilities after receiving feedback.

correction propagationevaluation metricsfeedback adaptation

This work addresses the prevalence of factual errors in deployed Retrieval-Augmented Generation (RAG) systems, which often stem from missing retrieval evidence or contextually inconsistent generation—issues that existing repair methods struggle to resolve under black-box conditions or resource constraints. To tackle this challenge, the authors propose D2R-RAG, a novel model-agnostic and resource-aware framework for diagnosing and repairing RAG failures. D2R-RAG employs lightweight modules to extract interpretable failure signatures from the query, retrieved passages, and generated response, then adaptively selects the optimal repair strategy under explicit latency and memory budgets. Experimental results demonstrate that D2R-RAG significantly enhances reliability on FEVER and HotpotQA, consistently outperforming existing baselines across diverse computational budgets while achieving a superior trade-off between accuracy and efficiency.

black-box settingbudget constraintsfactual errors