hallucination detection benchmarking

Designs and builds evaluation datasets, tasks, and metric suites that measure how well methods detect hallucinated content produced by vision-language and video-language models, covering temporal, motion, and token-level/localized hallucinations. Defines benchmarking protocols and constrained implementations (online, lightweight, CPU-/GPU-feasible), implements or integrates detectors (e.g., grounding, temporal cross-attention, vid-pair methods), and runs comparative analyses that report metrics (AUC, statistical significance) and organize results by model access level.

hallucinationdetectionbenchmarking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.54
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Current video large language models are prone to action hallucinations—generating descriptions inconsistent with visual content—due to co-occurrence priors, sequential reasoning errors, or confusion from fine-grained visual similarity. Existing evaluation benchmarks lack systematic coverage of such failure modes. This work proposes MoHallBench, the first fine-grained benchmark specifically designed to assess action hallucinations in videos, comprising 11,306 video clips and 40,493 question-answer pairs across binary-choice, multiple-choice, and generative tasks. To mitigate affirmation bias, it introduces a bidirectional questioning protocol and bias-aware metrics. Experiments on ten state-of-the-art models reveal that hallucinations stemming from sequential reasoning are most severe, and that strong priors or fine-grained similarity significantly exacerbate the issue. Notably, high action recognition accuracy does not guarantee low hallucination rates, indicating a decoupling between recognition capability and hallucination robustness.

action recognitionhallucination evaluationmotion hallucination

Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection

Apr 25, 2025
AK
Atharva Kulkarni
🏛️ University of Southern California | Apple Inc.

Reliable evaluation of language model hallucinations remains challenged by poor metric robustness, weak generalizability, and low agreement with human judgments. This work conducts the first large-scale empirical study to systematically assess 12 mainstream hallucination metrics across 4 datasets, 37 models, and 5 decoding strategies. Results reveal pervasive limitations: narrow evaluation scope, unstable gains under parameter scaling, and low inter-annotator agreement with human labels (mean Krippendorff’s α = 0.41). GPT-4 achieves the highest consistency as an evaluator (α = 0.72). Moreover, pattern-optimizing decoding—specifically Top-k and Nucleus sampling—reduces hallucination rates by 18.3% on average in knowledge-augmented settings. The study introduces a multidimensional benchmarking framework and advocates the LLM-as-a-judge paradigm, providing both empirical foundations and methodological guidance for building trustworthy hallucination evaluation systems.

Assessing robustness of hallucination detection metrics across diverse modelsEvaluating alignment between metrics and human judgments on hallucinationsIdentifying effective decoding methods to reduce language model hallucinations

ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models

Nov 16, 2024
VR
Vipula Rawte
🏛️ University of South Carolina | Guru Gobind Singh Indraprastha University | Vellore Institute of Technology | Indian Institute of Technology | University of Massachusetts | University of California | Amazon Web Services | Meta

This work addresses the pervasive hallucination problem in text-to-video (T2V) generation by large multimodal models (LMMs). We first systematically define and annotate five canonical hallucination categories, then construct ViBe—the first open-source, human-verified large-scale T2V hallucination benchmark—comprising 3,782 video samples generated from 837 COCO prompts. To detect hallucinations, we propose a video embedding framework combining TimeSformer and CNN features, coupled with ensemble classification, and establish a standardized human annotation protocol. Comprehensive evaluation across ten state-of-the-art T2V models reveals that the best-performing baseline achieves only 0.345 accuracy, underscoring the significant challenges in automated hallucination detection. The ViBe dataset and evaluation code are publicly released, providing critical infrastructure and a new standard for quantitatively assessing reliability and improving robustness of T2V models.

Developing benchmarks for reliable video generationEvaluating hallucination in Text-to-Video modelsIdentifying major hallucination types in generated videos

EventHallusion: Diagnosing Event Hallucinations in Video LLMs

Sep 25, 2024
JZ
Jiacheng Zhang
🏛️ Fudan University | Meituan

Video large language models (LLMs) suffer from severe hallucinations in event understanding, exacerbated by linguistic priors and vision-language misalignment. To address this, we introduce EventHallusion—the first benchmark dedicated to diagnosing event-level hallucinations in video LLMs—systematically defining, quantifying, and attributing such hallucinations from dual perspectives: linguistic priors and cross-modal biases. We propose temporal contrastive decoding (TCD), a novel, fine-tuning-free inference-time method that explicitly models temporal cues to suppress hallucination. Leveraging event-driven adversarial video construction and a multidimensional evaluation framework, we validate TCD across eight open-source and two closed-source models. Results show that TCD significantly enhances the reliability of event understanding, improving accuracy by over 15% for several models. This work establishes a new paradigm for trustworthy evaluation and robust reasoning in video LLMs.

Bias in Language and VisionMultimodal Large Language ModelsVisual Content Understanding

AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Oct 23, 2024
KS
Kim Sung-Bin
🏛️ POSTECH | KAIST | Yonsei University

Audio-visual large language models (AV-LLMs) suffer from cross-modal hallucination—erroneous associations between audio and visual signals—yet lack dedicated, standardized evaluation benchmarks. Method: We introduce AVHBench, the first benchmark specifically designed to evaluate cross-modal hallucination in AV-LLMs. It formally defines and quantifies such hallucinations, establishes a three-dimensional evaluation framework covering perception, alignment matching, and multimodal reasoning, and constructs a test set via a synergistic strategy combining multi-granularity aligned samples with human annotation and adversarial perturbation to enable fine-grained attribution analysis. Results: Experiments reveal that state-of-the-art AV-LLMs are consistently vulnerable to modality-crossing interference, inducing widespread hallucination. Crucially, fine-tuning solely on AVHBench significantly enhances hallucination robustness. This work provides foundational tools and a methodological framework for trustworthy evaluation and optimization of audio-visual multimodal models.

Address hallucinations caused by audio-visual signal misinterpretations.Develop a benchmark to improve robustness against hallucinations.Evaluate audio-visual LLMs' cross-modal perception and comprehension.

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing hallucination evaluation methods for vision-language models, which predominantly rely on model-generated samples and suffer from poor timeliness and insufficient controllability. To overcome these issues, the authors construct a dataset comprising 1,600 human-written, multilingual hallucination samples, complemented by fine-grained, span-level annotations. Through systematic comparative analysis, they demonstrate for the first time that human-authored samples significantly outperform model-generated ones in terms of distributional similarity, annotation consistency, and content controllability. These advantages enable more stable and generalizable assessment of model hallucination detection capabilities, offering a reliable alternative for benchmarking hallucinations in vision-language systems.

fine-grained annotationhallucination benchmarkinghuman-written samples

Existing evaluations of multimodal large language models predominantly focus on the manifestations of hallucinations rather than their underlying causes, often employing simplistic scenarios and limited assessment formats. This work proposes ReactBench, the first causality-driven benchmark grounded in the root causes of hallucinations. It constructs four diagnostic tasks—relation ablation, counterfactual attributes, image perturbation tracking, and dense counting—by leveraging adversarial images and suggestive queries, and integrates chain-of-thought reasoning to enable fine-grained attribution analysis. Designed in an exam-style evaluation format, ReactBench moves beyond conventional accuracy-only metrics. Experimental results demonstrate that current models remain highly susceptible to specific hallucination triggers, underscoring ReactBench’s effectiveness in diagnosing model weaknesses and enhancing robustness and interpretability.

benchmark limitationscause-driven evaluationmultimodal hallucination

Current multimodal large language models frequently suffer from motion hallucination in cross-video action comparison, struggling to accurately capture genuine kinematic differences. This work presents the first systematic characterization of motion hallucination along three dimensions—direction, attributes, and temporal structure—and introduces MotionHalluc, a benchmark comprising 553 video pairs and 1,540 fine-grained questions, to enable quantitative evaluation. The authors propose Perceive-Parse-Verify, a training-free method that translates natural language instructions into executable measurement queries and mitigates hallucination through explicit motion verification. Experiments reveal that motion hallucination is prevalent across mainstream models, while integrating measurement injection yields an average performance gain of 10.6%, underscoring the critical role of explicit motion measurement in improving model reliability.

cross-video comparisonfine-grained motionkinematic reasoning

This study addresses the potential overestimation of hallucination detection performance due to dataset construction artifacts—particularly prompt leakage—in existing benchmarks. Through a systematic evaluation of 22 detection methods, 12 open-source models, and 6 corpora, the work quantifies, for the first time, the extent to which such artifacts inflate reported results. To enable reliable real-time hallucination detection, the authors propose DRIFT, a supervised probing method based on transitions in upper-layer hidden states. Experimental findings reveal that, once prompt leakage is controlled, most existing approaches perform near random chance, with only SAPLMA and DRIFT demonstrating consistent effectiveness across diverse settings. These results indicate that current progress in hallucination detection has been substantially overstated and establish a more trustworthy evaluation framework for real-world applications.

benchmark artifactsevaluation biashallucination detection

Hot Scholars

BS

Bernt Schiele

Professor, Max Planck Institute for Informatics, Saarland University, Saarland Informatics Campus
Computer VisionMachine LearningArtificial IntelligenceAutonomous Driving
JW

Janet Wang

Tulane University
Medical AIVision Language ModelsGenerative AI
RC

Rajen Chatterjee

Apple Inc
Machine TranslationAutomatic Post-EditingNatural Language ProcessingCrowdsourcing
XC

Xilin Chen

Institute of Computing Technology, Chinese Academy of Sciences
Computer VisionPattern RecognitionMachine Learning