Score
Designs and builds evaluation datasets, tasks, and metric suites that measure how well methods detect hallucinated content produced by vision-language and video-language models, covering temporal, motion, and token-level/localized hallucinations. Defines benchmarking protocols and constrained implementations (online, lightweight, CPU-/GPU-feasible), implements or integrates detectors (e.g., grounding, temporal cross-attention, vid-pair methods), and runs comparative analyses that report metrics (AUC, statistical significance) and organize results by model access level.
Multimodal large language models (MLLMs) suffer from hallucination in image-to-text (I2T) and text-to-image (T2I) generation—producing outputs inconsistent with input images or real-world knowledge. This work systematically surveys hallucination phenomena across both tasks, proposing the first taxonomy that jointly characterizes fidelity (image alignment) and factuality (world-knowledge consistency). We unify and analyze existing evaluation benchmarks by distilling their construction principles and quantitative metrics, and categorize instance-level detection methods into three paradigms: output consistency analysis, external knowledge verification, and cross-modal feature inspection. Our analysis exposes critical limitations of prevailing datasets and methods in fine-grained hallucination localization, cross-modal alignment, and domain coverage. To address these gaps, we introduce a comprehensive, reliability-oriented evaluation framework for multimodal generation. This framework establishes a theoretical foundation and practical guidance for future research on hallucination mitigation and trustworthy multimodal AI.
Current video large language models are prone to action hallucinations—generating descriptions inconsistent with visual content—due to co-occurrence priors, sequential reasoning errors, or confusion from fine-grained visual similarity. Existing evaluation benchmarks lack systematic coverage of such failure modes. This work proposes MoHallBench, the first fine-grained benchmark specifically designed to assess action hallucinations in videos, comprising 11,306 video clips and 40,493 question-answer pairs across binary-choice, multiple-choice, and generative tasks. To mitigate affirmation bias, it introduces a bidirectional questioning protocol and bias-aware metrics. Experiments on ten state-of-the-art models reveal that hallucinations stemming from sequential reasoning are most severe, and that strong priors or fine-grained similarity significantly exacerbate the issue. Notably, high action recognition accuracy does not guarantee low hallucination rates, indicating a decoupling between recognition capability and hallucination robustness.
Reliable evaluation of language model hallucinations remains challenged by poor metric robustness, weak generalizability, and low agreement with human judgments. This work conducts the first large-scale empirical study to systematically assess 12 mainstream hallucination metrics across 4 datasets, 37 models, and 5 decoding strategies. Results reveal pervasive limitations: narrow evaluation scope, unstable gains under parameter scaling, and low inter-annotator agreement with human labels (mean Krippendorff’s α = 0.41). GPT-4 achieves the highest consistency as an evaluator (α = 0.72). Moreover, pattern-optimizing decoding—specifically Top-k and Nucleus sampling—reduces hallucination rates by 18.3% on average in knowledge-augmented settings. The study introduces a multidimensional benchmarking framework and advocates the LLM-as-a-judge paradigm, providing both empirical foundations and methodological guidance for building trustworthy hallucination evaluation systems.
This work addresses the pervasive hallucination problem in text-to-video (T2V) generation by large multimodal models (LMMs). We first systematically define and annotate five canonical hallucination categories, then construct ViBe—the first open-source, human-verified large-scale T2V hallucination benchmark—comprising 3,782 video samples generated from 837 COCO prompts. To detect hallucinations, we propose a video embedding framework combining TimeSformer and CNN features, coupled with ensemble classification, and establish a standardized human annotation protocol. Comprehensive evaluation across ten state-of-the-art T2V models reveals that the best-performing baseline achieves only 0.345 accuracy, underscoring the significant challenges in automated hallucination detection. The ViBe dataset and evaluation code are publicly released, providing critical infrastructure and a new standard for quantitatively assessing reliability and improving robustness of T2V models.
Video large language models (LLMs) suffer from severe hallucinations in event understanding, exacerbated by linguistic priors and vision-language misalignment. To address this, we introduce EventHallusion—the first benchmark dedicated to diagnosing event-level hallucinations in video LLMs—systematically defining, quantifying, and attributing such hallucinations from dual perspectives: linguistic priors and cross-modal biases. We propose temporal contrastive decoding (TCD), a novel, fine-tuning-free inference-time method that explicitly models temporal cues to suppress hallucination. Leveraging event-driven adversarial video construction and a multidimensional evaluation framework, we validate TCD across eight open-source and two closed-source models. Results show that TCD significantly enhances the reliability of event understanding, improving accuracy by over 15% for several models. This work establishes a new paradigm for trustworthy evaluation and robust reasoning in video LLMs.
Audio-visual large language models (AV-LLMs) suffer from cross-modal hallucination—erroneous associations between audio and visual signals—yet lack dedicated, standardized evaluation benchmarks. Method: We introduce AVHBench, the first benchmark specifically designed to evaluate cross-modal hallucination in AV-LLMs. It formally defines and quantifies such hallucinations, establishes a three-dimensional evaluation framework covering perception, alignment matching, and multimodal reasoning, and constructs a test set via a synergistic strategy combining multi-granularity aligned samples with human annotation and adversarial perturbation to enable fine-grained attribution analysis. Results: Experiments reveal that state-of-the-art AV-LLMs are consistently vulnerable to modality-crossing interference, inducing widespread hallucination. Crucially, fine-tuning solely on AVHBench significantly enhances hallucination robustness. This work provides foundational tools and a methodological framework for trustworthy evaluation and optimization of audio-visual multimodal models.
This work addresses the limitations of existing hallucination evaluation methods for vision-language models, which predominantly rely on model-generated samples and suffer from poor timeliness and insufficient controllability. To overcome these issues, the authors construct a dataset comprising 1,600 human-written, multilingual hallucination samples, complemented by fine-grained, span-level annotations. Through systematic comparative analysis, they demonstrate for the first time that human-authored samples significantly outperform model-generated ones in terms of distributional similarity, annotation consistency, and content controllability. These advantages enable more stable and generalizable assessment of model hallucination detection capabilities, offering a reliable alternative for benchmarking hallucinations in vision-language systems.
Existing evaluations of multimodal large language models predominantly focus on the manifestations of hallucinations rather than their underlying causes, often employing simplistic scenarios and limited assessment formats. This work proposes ReactBench, the first causality-driven benchmark grounded in the root causes of hallucinations. It constructs four diagnostic tasks—relation ablation, counterfactual attributes, image perturbation tracking, and dense counting—by leveraging adversarial images and suggestive queries, and integrates chain-of-thought reasoning to enable fine-grained attribution analysis. Designed in an exam-style evaluation format, ReactBench moves beyond conventional accuracy-only metrics. Experimental results demonstrate that current models remain highly susceptible to specific hallucination triggers, underscoring ReactBench’s effectiveness in diagnosing model weaknesses and enhancing robustness and interpretability.
Current multimodal large language models frequently suffer from motion hallucination in cross-video action comparison, struggling to accurately capture genuine kinematic differences. This work presents the first systematic characterization of motion hallucination along three dimensions—direction, attributes, and temporal structure—and introduces MotionHalluc, a benchmark comprising 553 video pairs and 1,540 fine-grained questions, to enable quantitative evaluation. The authors propose Perceive-Parse-Verify, a training-free method that translates natural language instructions into executable measurement queries and mitigates hallucination through explicit motion verification. Experiments reveal that motion hallucination is prevalent across mainstream models, while integrating measurement injection yields an average performance gain of 10.6%, underscoring the critical role of explicit motion measurement in improving model reliability.
This study addresses the potential overestimation of hallucination detection performance due to dataset construction artifacts—particularly prompt leakage—in existing benchmarks. Through a systematic evaluation of 22 detection methods, 12 open-source models, and 6 corpora, the work quantifies, for the first time, the extent to which such artifacts inflate reported results. To enable reliable real-time hallucination detection, the authors propose DRIFT, a supervised probing method based on transitions in upper-layer hidden states. Experimental findings reveal that, once prompt leakage is controlled, most existing approaches perform near random chance, with only SAPLMA and DRIFT demonstrating consistent effectiveness across diverse settings. These results indicate that current progress in hallucination detection has been substantially overstated and establish a more trustworthy evaluation framework for real-world applications.