🤖 AI Summary
This work addresses the limitations of traditional retrieval-augmented generation (RAG) systems, which rely on flat knowledge snippets and single-step embedding-based retrieval, thereby struggling to support multi-hop reasoning, evidence evaluation, and structured traversal required by intelligent agents. The authors propose LLM-Wiki, a novel framework that reconceptualizes retrieval as an integral part of reasoning by constructing a composable, evolvable, and bidirectionally linked structured Wiki knowledge base. Through a standardized tool-calling interface, the system enables search, reading, and link-tracing capabilities. It further incorporates an error-log-driven self-correction mechanism that synergistically combines large language models with knowledge graph techniques to continuously refine both knowledge structure and semantic representation. Evaluated on HotpotQA, MuSiQue, and 2WikiMultiHopQA, LLM-Wiki outperforms seven state-of-the-art baselines—including HippoRAG 2, LightRAG, and GraphRAG—achieving F1 score improvements of 2.0–8.1 points, with particularly strong performance on the AuthTrace multi-document structured query task.
📝 Abstract
LLM agents require retrieval to behave less like one-shot context fetching and more like reasoning: searching, reading, traversing, and deciding when evidence is sufficient. However, Retrieval-Augmented Generation (RAG) typically organizes external knowledge as flat chunks retrieved by embedding similarity, exposing a retrieval-as-lookup interface that is poorly aligned with tool-using agents. We propose LLM-Wiki, an agent-native retrieval system that operationalizes the Retrieval-as-Reasoning paradigm by treating external knowledge as a compilable, composable, and self-evolving structure rather than a static retrieval index. LLM-Wiki compiles documents into structured Wiki pages with bidirectional links, exposes search, read, and link-following operations through standard tool-calling interfaces, and introduces an Error Book for persistent structural and semantic self-correction. On HotpotQA, MuSiQue, and 2WikiMultiHopQA, LLM-Wiki outperforms seven baselines, including HippoRAG 2, LightRAG, and GraphRAG, with gains of 2.0-8.1 F1 points over the strongest graph-based baseline and larger gains over Dense RAG. On AuthTrace, LLM-Wiki achieves the best overall accuracy, with especially strong gains on multi-document structured queries, showing that compilation-based knowledge organization generalizes beyond chain-style multi-hop reasoning.