🤖 AI Summary
This study addresses the limitation of dense retrieval systems that rely on surface-level semantics and struggle with complex search tasks. To overcome this bottleneck, we propose a ReAct agent architecture that integrates large language model reasoning with embedding-based retrieval. By constructing a "reasoning-acting" loop for multi-step retrieval, this approach transcends traditional paradigms and achieves strong cross-domain generalization. Experimental results demonstrate that the proposed system yields competitive performance on the ViDoRe v3 and BRIGHT benchmarks, improving nDCG@10 by 8.7 points. However, the average query latency reaches 107.4 seconds alongside substantial token consumption, revealing a fundamental trade-off between retrieval effectiveness and computational cost.
📝 Abstract
Modern information systems, including many agentic workflows, use dense retrieval to explore large amounts of unstructured data. However, dense retrieval relies on surface-level semantic similarity, which is insufficient for increasingly complex search applications. Here, we investigate agentic retrieval that combines the reasoning capabilities of Large Language Models (LLMs) with the efficient corpus exploration of retrievers in a ReAct agentic loop to solve complex retrieval tasks. In our experiments, we show that agentic retrieval is more effective than standard retrieval, improving nDCG@10 by 8.7 points using the same embedding model. Moreover, while specialized retrieval methods struggle on out-of-domain tasks, agentic retrieval is highly generalizable: the same pipeline achieves competitive results on both the ViDoRe v3 and BRIGHT leaderboards. However, this improvement comes at a cost. On average, agentic retrieval takes 107.4 seconds, compared to 0.67 seconds for standard retrieval, and consumes 764.1K input and 5.8K output tokens per query. In short, our study demonstrates the effectiveness of agentic retrieval in modern data systems and motivates future work on more cost-efficient retrieval agents for large-scale deployment.