🤖 AI Summary
This work addresses the susceptibility of large language models to factual hallucinations in complex question answering. To mitigate this issue, the authors propose a lightweight graph-augmented retrieval-augmented generation system that integrates vector retrieval with graph-based query tools, enabling efficient multi-hop reasoning over a Wikipedia subset. By incorporating a concise graph schema and a dedicated toolset, the approach effectively curbs hallucinations and enhances fine-grained factual accuracy without substantially increasing computational overhead. Experimental results demonstrate a 50% reduction in hallucinated responses compared to baseline methods, along with significant improvements in both precision and recall for factual correctness. The method achieves state-of-the-art performance on the MoNaCo benchmark in terms of fine-grained truthfulness.
📝 Abstract
Large language models (LLMs) have fundamentally transformed the landscape of Natural Language Processing. Despite these advances, LLMs and LLM-based systems remain prone to a variety of failure modes. Retrieval-augmented generation (RAG) systems have emerged as a common deployment scenario seeking to both avoid the well known risk of the LLM "hallucinating" information, and to enable reasoning and question answering over proprietary information that the LLM did not have access to during training without resorting to expensive model fine-tuning.
In this work, we explore the idea of using a lightweight graph structure with a relatively simple graph schema, to support the RAG subsystem via a dedicated toolset. We design an agentic system with a variety of vector search and graph query tools operating over a structured dataset based on a curated subset of English Wikipedia articles, and evaluate its performance on questions from MoNaCo, a challenging Wikipedia QA benchmark of complex query answering tasks.
Our results show that the introduction of graph-based tools can significantly increase the precision and recall of factual correctness, can halve the number of hallucinated answers, and achieves the highest fine-grained truthfulness score among the three evaluated scenarios. All this with a modest increase in token usage.