🤖 AI Summary
In enterprise RAG scenarios, large language models (LLMs) struggle with domain-ambiguous queries and often introduce lexical or semantic noise. To address this, we propose VERDICT—a novel framework featuring a Verified-Diversification mechanism that enables collaborative, feedback-driven interaction between retriever and generator. Unlike conventional cascaded “generate–retrieve–filter” pipelines, VERDICT integrates differentiable early-stage verification with diverse explanation generation in a closed loop, mitigating error propagation. Our method unifies dynamic explanation generation, retrieval feedback modeling, and consistency consolidation. Evaluated on the ASQA benchmark, VERDICT achieves an average 23% improvement in grounded F1 over strong baselines. Moreover, it demonstrates robust generalization across multiple mainstream LLM backbones. VERDICT establishes a new paradigm for ambiguity resolution in RAG—robust, efficient, and scalable—while preserving interpretability and grounding fidelity.
📝 Abstract
In this work, we tackle the challenge of disambiguating queries in retrieval-augmented generation (RAG) to diverse yet answerable interpretations. State-of-the-arts follow a Diversify-then-Verify (DtV) pipeline, where diverse interpretations are generated by an LLM, later used as search queries to retrieve supporting passages. Such a process may introduce noise in either interpretations or retrieval, particularly in enterprise settings, where LLMs -- trained on static data -- may struggle with domain-specific disambiguations. Thus, a post-hoc verification phase is introduced to prune noises. Our distinction is to unify diversification with verification by incorporating feedback from retriever and generator early on. This joint approach improves both efficiency and robustness by reducing reliance on multiple retrieval and inference steps, which are susceptible to cascading errors. We validate the efficiency and effectiveness of our method, Verified-Diversification with Consolidation (VERDICT), on the widely adopted ASQA benchmark to achieve diverse yet verifiable interpretations. Empirical results show that VERDICT improves grounding-aware F1 score by an average of 23% over the strongest baseline across different backbone LLMs.