🤖 AI Summary
This study addresses the limitation that hallucination detection in retrieval-augmented generation (RAG) typically depends on specific generators, hindering adaptation to closed-source or alternative models. To overcome this, we propose a generator-decoupled hallucination detection framework. This method is the first to leverage the hidden states of an independent observer model to construct grounding probes, eliminating reliance on internal generator activations through a single forward pass. Specifically, it applies logistic regression to pooled intermediate-layer features and integrates these with a supervised span detector for final prediction. Evaluated on the RAGTruth benchmark, the proposed framework achieves AUROC scores of 0.879–0.894 and maintains robust performance across six distinct generators, significantly outperforming existing baseline methods.
📝 Abstract
Detecting responses that retrieval-augmented generation does not ground in its context trades speed against accuracy: surface checks miss paraphrased fabrication, sampling-based methods cost extra generations. Hidden-state probes sit between the two, but every existing one reads the generating model's own activations, so a change of generator invalidates the detector and a closed-weight generator is out of reach. This paper removes that coupling. The Grounding Probe is logistic regression over the mean-pooled middle-layer hidden states of an observer language model that reads the context, question, and response in one forward pass and generates nothing, with the recipe it needs: pool over response tokens, read a middle layer, and control capacity, which closes the train-test AUROC gap from 0.087-0.202 to 0.009-0.013. Asking the observer outright, rather than reading its hidden state, costs at least +0.166 AUROC in every one of four models. Fitted on 15,090 annotated responses it reaches 0.879-0.894 AUROC on RAGTruth test across four observers, and 0.924 AUROC with 0.820 F1@0.5 averaged with a supervised span detector, 0.060 above that detector alone. One probe holds across six generators, and hold-out controls, including one in which no evaluation prompt appears in training, bound the cost of removing a generator at about 0.02 AUROC. Code, probes, and predictions are released.