RIT-RAG: Navigating Document Corpora with Retrieval-Induced Trees

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of Agentic RAG in lacking document structure awareness and the difficulty of scaling traditional methods to large corpora by proposing a Retrieval-Induced Tree mechanism. Specifically, this method constructs document trees offline and employs an online subtree induction algorithm to generate traversable structured subtrees for agents, thereby integrating content retrieval with structural navigation to enable broad cross-document retrieval alongside deep comprehension. Evaluated across multi-domain benchmarks, the proposed approach achieves state-of-the-art accuracy, outperforming the strongest baseline by 6.8 to 11.4 percentage points on EntQABench.
📝 Abstract
Retrieval-augmented generation (RAG) grounds language models in external corpora. Agentic RAG enables iterative search, yet exposes the model to isolated chunks without document structure, making it difficult to distinguish relevant evidence from chunks that merely resemble the query. Structure-aware methods such as PageIndex navigate document structure but cannot scale to the structures of large corpora, which do not fit in the LLM context. Hence, they first commit to a single document using a document retriever and cannot recover from a wrong choice. We propose RIT-RAG (Retrieval-Induced Tree RAG), which combines content retrieval with structural navigation. Offline, RIT-RAG builds a tree for each document from its table of contents or sitemap. At query time, it retrieves a broad set of chunks and uses their positions to induce manageable sub-trees, potentially across multiple documents. An LLM agent navigates these sub-trees, selectively reads promising nodes, and reformulates queries when needed. Thus, retrieval proposes where to look, while the agent decides what to read. Across financial, scientific, and customer-support benchmarks, RIT-RAG achieves the highest answer accuracy among vanilla, graph-based, and agentic baselines. On EntQABench, our new benchmark of 2.84 million technical-documentation webpages, it improves accuracy by 6.8 to 11.4 points over the strongest baseline across three LLMs.
Problem

Research questions and friction points this paper is trying to address.

Retrieval-Augmented Generation
Document Structure
Large Corpora
Agentic RAG
Structural Navigation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Retrieval-Augmented Generation
Retrieval-Induced Trees
Agentic RAG
Structure-aware Navigation
Document Corpora
🔎 Similar Papers
No similar papers found.