VOPE: Revisiting Hallucination of Vision-Language Models in Voluntary Imagination Task

📅 2025-11-17
📈 Citations: 0
Influential: 0
📄 PDF

career value

186K/year
🤖 AI Summary
Existing hallucination studies on large vision-language models (LVLMs) focus predominantly on factual description tasks, overlooking the fundamental distinction between “reasonable creativity” and “inappropriate hallucination” in voluntary imagination tasks—such as image-based story generation. This work introduces VOPE, a novel hallucination evaluation framework that explicitly distinguishes these two task categories. VOPE employs a re-examination questioning mechanism to assess model self-consistency in judging the existence of its own generated content, jointly grounded in the actual visual content of the input image. By decoupling task-aligned creativity from spurious generation, VOPE avoids misclassifying legitimate creative outputs as hallucinations. Empirical evaluation reveals pervasive and severe hallucinations in story generation across mainstream LVLMs, with current mitigation methods demonstrating limited efficacy. VOPE establishes a new paradigm and an empirically validated benchmark for hallucination assessment in open-ended generative scenarios.

Technology Category

Application Category

📝 Abstract
Most research on hallucinations in Large Vision-Language Models (LVLMs) focuses on factual description tasks that prohibit any output absent from the image. However, little attention has been paid to hallucinations in voluntary imagination tasks, e.g., story writing, where the models are expected to generate novel content beyond the given image. In these tasks, it is inappropriate to simply regard such imagined novel content as hallucinations. To address this limitation, we introduce Voluntary-imagined Object Presence Evaluation (VOPE)-a novel method to assess LVLMs' hallucinations in voluntary imagination tasks via presence evaluation. Specifically, VOPE poses recheck-based questions to evaluate how an LVLM interprets the presence of the imagined objects in its own response. The consistency between the model's interpretation and the object's presence in the image is then used to determine whether the model hallucinates when generating the response. We apply VOPE to several mainstream LVLMs and hallucination mitigation methods, revealing two key findings: (1) most LVLMs hallucinate heavily during voluntary imagination, and their performance in presence evaluation is notably poor on imagined objects; (2) existing hallucination mitigation methods show limited effect in voluntary imagination tasks, making this an important direction for future research.
Problem

Research questions and friction points this paper is trying to address.

Evaluating hallucinations in vision-language models during voluntary imagination tasks
Assessing model consistency in interpreting self-generated imagined content
Analyzing limitations of existing hallucination mitigation methods in creative tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluates hallucinations via recheck-based presence questions
Measures consistency between model interpretation and image content
Assesses voluntary imagination tasks in vision-language models
🔎 Similar Papers
No similar papers found.