Score
Designs and implements agentic retrieval-augmented generation (RAG) pipelines and LangChain-style agent orchestrations that retrieve multi-source contextual evidence, invoke tools, and apply LLM reasoning to convert evidence into structured JSON diagnoses and prioritized, actionable recommendations. Builds and analyzes components for iterative diagnosis, tool selection and orchestration, and mechanisms to incorporate operator feedback into reflective memory to refine subsequent reasoning and outputs.
To address the limitations of existing RAG systems in complex reasoning, dynamic retrieval, and multimodal integration within real-world industrial applications, this paper proposes an inference-enhanced intelligent RAG framework. Methodologically, it introduces the first dual-track reasoning taxonomy—System 1 (fast, modular reasoning) and System 2 (slow, autonomous planning)—and establishes the first open-source knowledge-graph-based RAG survey repository. The framework integrates LLM-driven reasoning architectures, standardized tool-use protocols (e.g., ReAct), multi-stage retrieval strategies, and multimodal interfaces. Through a systematic analysis of over 120 state-of-the-art works, we identify seven inference patterns and five solutions to key industrial bottlenecks. Empirical evaluation in production scenarios—including customer service and financial risk control—demonstrates 23%–38% improvements in reasoning accuracy.
Traditional RAG systems suffer from limited flexibility and contextual adaptability in complex, real-time, and multi-domain tasks. To address this, we propose Agentic RAG—a paradigm centered on autonomous AI agents that transcend static retrieval by enabling dynamic planning, reflective iteration, tool invocation, and multi-agent collaboration. Our contributions are threefold: (1) the first unified taxonomy for Agentic RAG; (2) a novel agent-driven mechanism for dynamic retrieval control; and (3) a closed-loop, task-adaptive workflow supporting multi-step reasoning. The approach integrates large language models, knowledge graphs, vector databases, and agent architectures. Extensive experiments span healthcare, finance, and education domains. We further provide scalable system design, ethics-aligned governance strategies, performance optimization guidelines, and an open-source engineering deployment framework.
Current agentic RAG systems lack a unified theoretical framework, resulting in architectural fragmentation, inconsistent evaluation protocols, and insufficiently characterized reliability risks. This work addresses these challenges by formally modeling agentic RAG as a finite-horizon partially observable Markov decision process. It introduces a modular architecture and a systematic taxonomy encompassing core components such as planning, retrieval coordination, memory paradigms, and tool invocation. By analyzing the limitations of static evaluation and identifying dynamic risks inherent in autonomous loops—particularly hallucination propagation and memory contamination—the study establishes a theoretical foundation for agentic RAG and proposes reliability-oriented evaluation criteria. Furthermore, it outlines key directions for future research, including adaptive retrieval strategies, cost-aware coordination mechanisms, and effective supervision frameworks.
This study addresses critical limitations of traditional Retrieval-Augmented Generation (RAG) systems—such as retrieval noise, misuse of retrieved content, weak query-document alignment, and high generation costs—and presents the first large-scale empirical comparison between enhanced RAG and agentic RAG paradigms. Leveraging a large language model–driven agent control flow, modular RAG components, and a multidimensional evaluation framework encompassing accuracy, robustness, and computational cost, the work systematically assesses performance and efficiency across diverse scenarios. Findings reveal that agentic RAG demonstrates superior adaptability in complex tasks, whereas enhanced RAG achieves higher efficiency in simpler settings. These results provide clear guidance on the trade-offs between performance and cost for real-world deployment and underscore a pathway toward more autonomous, agent-based RAG architectures.
Enterprise-grade RAG systems in high-stakes decision-making are often hindered by shallow retrieval, lack of traceability, and fragility to ambiguous queries. To address these limitations, this work proposes ADORE, a framework that orchestrates multiple specialized agents through a central coordinator to perform user-guided, iterative deep retrieval and synthesis. Key innovations include a structured memory repository based on Claim-Evidence Graphs, a memory-locking synthesis mechanism, an evidence-coverage-guided execution pipeline, segmented packing and compression for long-context handling, and an evidence-driven termination criterion—collectively enabling the generation of verifiable, fully traceable reports. ADORE achieves state-of-the-art performance with a score of 52.65 on the DeepResearch Bench and outperforms existing commercial systems by a 77.2% preference win rate in the DeepConsult evaluation.
This work addresses the overreliance of existing agent-based RAG systems on complex retrieval backends, which constrains the ability of large language models (LLMs) to precisely articulate information needs. The authors propose a novel agent RAG framework that restores retrieval control to the LLM by enabling it to explicitly formulate query intent through generated logical expressions, coupled with a lightweight inverted index for structured retrieval. By eschewing sophisticated embedding mechanisms, the approach substantially simplifies system architecture while maintaining performance comparable to strong baselines. This design significantly reduces both construction and serving costs and effectively mitigates hallucination in generated outputs.
This work addresses the limitations of conventional retrieval-augmented generation (RAG) pipelines when retrieval contexts are noisy, incomplete, or contain evidence scattered across multiple heterogeneous documents. To overcome these challenges, the authors propose a multi-agent collaborative RAG framework in which specialized agents perform distinct roles—evidence summarization, key information extraction, and multi-step reasoning—and subsequently synthesize their perspectives to generate a coherent final answer. Leveraging large language models to orchestrate this collaborative agent system, the approach demonstrates significant performance gains over strong baselines across four benchmark datasets, with particularly pronounced improvements in scenarios where relevant evidence is dispersed. These results substantiate the efficacy of multi-perspective analysis and collaborative synthesis in enhancing RAG robustness and accuracy.
Current retrieval-augmented generation (RAG) systems typically employ fixed retrieval pipelines, which struggle to accommodate the diverse demands of tasks such as factual question answering, multi-hop reasoning, and scientific claim verification. This work proposes Experience-RAG, an agent-oriented, plug-and-play skill module situated between the agent and a pool of multiple retrievers. It dynamically selects the optimal retrieval strategy by querying an experience memory and returns structured evidence. The key innovation lies in encapsulating retrieval strategy selection as a reusable agent skill, integrating experience-driven policy scheduling with a multi-retriever collaboration mechanism to enable flexible and adaptive retrieval orchestration. Experiments demonstrate that the method achieves an nDCG@10 of 0.8924 on BeIR/nq, BeIR/hotpotqa, and BeIR/scifact, significantly outperforming fixed-retrieval baselines and matching the performance of state-of-the-art adaptive RAG routing approaches.
In financial technology domains, dense terminology and pervasive acronyms degrade retrieval and generation performance in Retrieval-Augmented Generation (RAG) systems. To address this, we propose a modular multi-agent collaborative RAG framework that integrates context-aware acronym resolution, keyword-guided iterative subquery decomposition, intelligent query rewriting, and cross-encoder re-ranking—enabling end-to-end precise retrieval and generation of domain-specific knowledge. Evaluated on an enterprise-scale financial knowledge base across 85 expert-curated question-answer pairs, our framework significantly outperforms conventional RAG baselines: retrieval accuracy improves by 23.6%, and relevance scores increase by 19.4%. These results validate the efficacy of the multi-agent collaboration paradigm in highly specialized domains. Although introducing moderate latency, the substantial gains in accuracy demonstrably outweigh the computational overhead.
This work addresses the limitations of traditional retrieval-augmented generation (RAG) in handling complex queries—such as multi-hop reasoning and structured knowledge acquisition—where static, single-step retrieval proves inadequate. The authors propose an agent-driven adaptive RAG framework that dynamically decomposes queries, performs iterative retrieval, and incorporates a lightweight self-reflection evaluation loop to adjust retrieval strategies on demand. For the first time, the study systematically compares the efficacy of query decomposition and reflection mechanisms in structured and multi-hop settings, revealing that agent augmentation is not universally beneficial and advocating for cost-sensitive, adaptive orchestration. Experiments show a 0.04 improvement in overall score and a 0.17 gain in MRR on the DevOps dataset; however, query decomposition reduces ranking accuracy in multi-hop tasks, and while reflection enhances citation precision, it introduces notable latency.
Existing agent-based RAG approaches improve LLM reliability via reinforcement learning but incur substantial token overhead in retrieval and reasoning, trading efficiency for accuracy. This paper proposes an efficient RAG framework addressing this trade-off. First, it introduces a retrieval compression mechanism integrating knowledge-association graphs with personalized PageRank, enabling joint semantic chunk retrieval, graph-structured triplet retrieval, and knowledge matching. Second, it proposes Iterative Process-aware Direct Preference Optimization (IP-DPO), which explicitly models and optimizes the number of reasoning steps. Evaluated on six benchmarks, our method improves accuracy by 4% (Llama3-8B) and 2% (Qwen2.5-14B) on average, while reducing output tokens by 61% and 59%, respectively—demonstrating significant gains in both accuracy and generation efficiency.
This study addresses the challenge of achieving high-precision, interpretable question answering in specialized technical domains such as intelligent tires and vehicle dynamics. To this end, the authors propose a 13-step autonomous Retrieval-Augmented Generation (RAG) framework that integrates intent classification, multidimensional evidence sufficiency scoring, path-dependent external retrieval, and a post-generation self-correction loop. The framework further incorporates large language model (LLM)- and citation-driven knowledge graph construction using Neo4j. By synthesizing multi-source scholarly data from Crossref, OpenAlex, and Semantic Scholar, and employing a hybrid review mechanism combining rule-based and LLM-based evaluation alongside citation completeness verification, the approach significantly enhances retrieval relevance, answer accuracy, and citation reliability across a corpus of 2,100 research papers, enabling robust reasoning for complex technical queries.
This work addresses the performance limitations of existing dynamic agent-based RAG systems, which stem from the mismatch between planning strategies and execution capabilities due to their independent optimization. To overcome this, we propose JADE, a novel framework that models dynamic multi-turn RAG as a collaborative multi-agent team sharing a unified backbone. By leveraging end-to-end reinforcement learning with outcome-based rewards, JADE enables joint optimization of planning and execution. The approach introduces dynamic workflow orchestration, effectively bridging the gap between strategic intent and operational capacity. This integration not only maintains computational efficiency but also significantly enhances overall system performance through synergistic coordination and flexible trade-offs among modules.