Agentic AutoRAG: RAG Pipeline Optimization through Reasoning-Driven Agents

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive optimization costs and lack of failure attribution associated with tuning multiple hyperparameters in retrieval-augmented generation (RAG) pipelines. We propose an automated configuration framework driven by large language model (LLM) agents, which introduces a fine-grained failure attribution mechanism for both retrieval and generation stages. Specifically, a diagnostic agent and a proposal agent collaborate synergistically, leveraging knowledge-enhanced decision-making to construct a multi-objective Pareto optimization framework that balances accuracy and computational cost. Experimental results demonstrate that our approach comprehensively outperforms existing baselines across three benchmarks, achieving the performance of 30-trial statistical search methods within merely 10 trials. Furthermore, in medical domain scenarios, it attains 77% accuracy at only 58% of the standard cost, substantially improving the configuration efficiency and cost-effectiveness of RAG systems.
📝 Abstract
Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacting choices, from chunking and embedding model to reranking and generation. Existing optimizers, from greedy search to Bayesian optimization, reduce each trial to an aggregate score and search without modeling why a configuration performed as it did, even though the retrieved chunks already provide evidence about whether each failure occurred during retrieval or after it. We introduce Agentic AutoRAG, an LLM-agent optimizer for multi-objective RAG hyperparameter optimization with retrieval-versus-generation failure attribution. It proposes configurations scored on a frozen exam from the corpus: after each trial a Diagnoser attributes each failed question to retrieval or generation, and a Proposer, grounded in a knowledge base of model rankings and pricing, selects the next configuration, weighing accuracy against cost to trace a Pareto frontier. On three multi-hop QA benchmarks it reaches higher LLM-judge accuracy than every baseline we compare, and within its first 10 trials it matches or beats the statistical baselines' full 30-trial judge accuracy. In its cost-aware mode on a real-world healthcare corpus it reaches a median exam accuracy of 77%, above the strongest baseline's 71.5%, at about 58% of that baseline's cost per query, and it matches that 71.5% at about 22% of the cost.
Problem

Research questions and friction points this paper is trying to address.

Retrieval-Augmented Generation
Hyperparameter Optimization
Failure Attribution
Multi-Objective Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Retrieval-Augmented Generation
Agentic Optimization
Failure Attribution
Multi-Objective Hyperparameter Tuning
Pareto Frontier
🔎 Similar Papers
No similar papers found.
L
Lasse B. Strand
ETH Zurich
R
Robert Jakob
ETH Zurich
K
Kevin O'Sullivan
ETH Zurich
Markus Kreft
Markus Kreft
ETH Zurich
machine learningenergy efficiencysmart gridelectric vehiclessustainability