Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking

📅 2026-08-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing hallucination evaluation methods for vision-language models, which predominantly rely on model-generated samples and suffer from poor timeliness and insufficient controllability. To overcome these issues, the authors construct a dataset comprising 1,600 human-written, multilingual hallucination samples, complemented by fine-grained, span-level annotations. Through systematic comparative analysis, they demonstrate for the first time that human-authored samples significantly outperform model-generated ones in terms of distributional similarity, annotation consistency, and content controllability. These advantages enable more stable and generalizable assessment of model hallucination detection capabilities, offering a reliable alternative for benchmarking hallucinations in vision-language systems.
📝 Abstract
In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-generated hallucinations, in order to make benchmarking detection independent of particular models. To this end, we construct a dataset of 1,600 human-written samples, spanning four languages (Chinese, English, French, Italian), and 18,400 samples from five vision-and-language models, all annotated for hallucinations using a fine-grained span-level labeling scheme. We find that human-written samples result in higher agreement and allow greater control of dataset contents, while remaining distributionally similar to samples derived from vision-and-language samples and providing a reasonable portrayal of detection capabilities - suggesting that human data is a viable substitute for model-based hallucination benchmarks.
Problem

Research questions and friction points this paper is trying to address.

hallucination benchmarking
vision-and-language models
human-written samples
model evaluation
fine-grained annotation
Innovation

Methods, ideas, or system contributions that make the work stand out.

human-written hallucination
fine-grained benchmarking
vision-and-language models
span-level annotation
model-independent evaluation
🔎 Similar Papers