🤖 AI Summary
This study addresses the prohibitive optimization costs and lack of failure attribution associated with tuning multiple hyperparameters in retrieval-augmented generation (RAG) pipelines. We propose an automated configuration framework driven by large language model (LLM) agents, which introduces a fine-grained failure attribution mechanism for both retrieval and generation stages. Specifically, a diagnostic agent and a proposal agent collaborate synergistically, leveraging knowledge-enhanced decision-making to construct a multi-objective Pareto optimization framework that balances accuracy and computational cost. Experimental results demonstrate that our approach comprehensively outperforms existing baselines across three benchmarks, achieving the performance of 30-trial statistical search methods within merely 10 trials. Furthermore, in medical domain scenarios, it attains 77% accuracy at only 58% of the standard cost, substantially improving the configuration efficiency and cost-effectiveness of RAG systems.
📝 Abstract
Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacting choices, from chunking and embedding model to reranking and generation. Existing optimizers, from greedy search to Bayesian optimization, reduce each trial to an aggregate score and search without modeling why a configuration performed as it did, even though the retrieved chunks already provide evidence about whether each failure occurred during retrieval or after it. We introduce Agentic AutoRAG, an LLM-agent optimizer for multi-objective RAG hyperparameter optimization with retrieval-versus-generation failure attribution. It proposes configurations scored on a frozen exam from the corpus: after each trial a Diagnoser attributes each failed question to retrieval or generation, and a Proposer, grounded in a knowledge base of model rankings and pricing, selects the next configuration, weighing accuracy against cost to trace a Pareto frontier. On three multi-hop QA benchmarks it reaches higher LLM-judge accuracy than every baseline we compare, and within its first 10 trials it matches or beats the statistical baselines' full 30-trial judge accuracy. In its cost-aware mode on a real-world healthcare corpus it reaches a median exam accuracy of 77%, above the strongest baseline's 71.5%, at about 58% of that baseline's cost per query, and it matches that 71.5% at about 22% of the cost.