đ€ AI Summary
Retrieval-Augmented Generation (RAG) systems exhibit insufficient reliability in high-stakes domainsâsuch as investment due diligenceâwhere hallucinations, off-topic responses, and citation failures pose critical risks.
Method: We propose a statistically rigorous yet industrially scalable evaluation framework that integrates human expert annotations with calibrated LLM-based judges in a collaborative assessment protocol. Inspired by predictive inference, our approach enables precise, quantitative identification of multiple failure modesâincluding hallucination, irrelevance, and retrieval failure.
Contributions/Results: (1) We introduce the first RAG reliability benchmark specifically designed for fund due diligence; (2) we publicly release a high-quality evaluation dataset and open-source evaluation code; (3) empirical validation demonstrates that our framework significantly improves assessment reliability while enabling large-scale automated evaluationâestablishing a reproducible, verifiable evaluation paradigm for deploying RAG systems in mission-critical applications.
đ Abstract
The rise of generative AI, has driven significant advancements in high-risk sectors like healthcare and finance. The Retrieval-Augmented Generation (RAG) architecture, combining language models (LLMs) with search engines, is particularly notable for its ability to generate responses from document corpora. Despite its potential, the reliability of RAG systems in critical contexts remains a concern, with issues such as hallucinations persisting. This study evaluates a RAG system used in due diligence for an investment fund. We propose a robust evaluation protocol combining human annotations and LLM-Judge annotations to identify system failures, like hallucinations, off-topic, failed citations, and abstentions. Inspired by the Prediction Powered Inference (PPI) method, we achieve precise performance measurements with statistical guarantees. We provide a comprehensive dataset for further analysis. Our contributions aim to enhance the reliability and scalability of RAG systems evaluation protocols in industrial applications.