🤖 AI Summary
This study addresses the challenge of detecting hallucinations in large language models under realistic black-box settings involving only a single query and a question-answer pair. The authors propose an operational hallucination classification method grounded in the geometric structure of embedding spaces, establishing—for the first time—a connection between hallucination detectability and the unit hyperspherical embeddings induced by contrastive learning. The approach diagnoses hallucinations by analyzing the directional relationship between the model’s response vector and regions corresponding to trustworthy answers. Integrating sentence encoder embeddings, angular ratios, and directional features, the method is benchmarked against natural language inference (NLI) baselines. Evaluation on a newly curated dataset of 212 human-crafted hallucinated QA pairs spanning nine domains demonstrates that directional features outperform NLI in identifying expert-annotated errors, while also revealing that factual errors sharing lexical or framing similarities cannot be reliably distinguished using angular geometry alone.
📝 Abstract
Hallucinations in deployed language models can have real consequences for downstream decisions in domains such as healthcare, legal, and financial services. In production, detection has to run on what the deployed system can see: the query, the response, and often a source document. White-box access to model internals and multi-sample querying are not generally available behind a third-party API. Within this setting - black-box, single-pass, only question/answer available - the dominant baseline is NLI, which returns a value but no diagnosis when it fails. We argue that operating directly on the geometry of the embedding space provides detection methods whose successes and failures are interpretable as structural properties of contrastive sentence-encoder training \citep{wang2020understanding}. The contribution is: given an operationally-motivated taxonomy, geometry predicts which types of hallucination are detectable and which are not - and the predictions hold. We propose three operational types organized by the relation of the response embedding to the plausibility region of grounded responses on the unit hypersphere, and derive from the alignment objective a prediction for each: (1)query-proximate unfaithfulness is detectable by an angular ratio; (2)confabulation outside the plausibility region produces a directional signature that outperforms NLI on expert-annotated error; (3)factual errors sharing vocabulary and frame with correct answers are not separable by angular geometry. To validate on content resembling deployment, we built a 212-pair human-confabulated dataset across nine domains using provoked confabulation.