🤖 AI Summary
Current video large language models are prone to action hallucinations—generating descriptions inconsistent with visual content—due to co-occurrence priors, sequential reasoning errors, or confusion from fine-grained visual similarity. Existing evaluation benchmarks lack systematic coverage of such failure modes. This work proposes MoHallBench, the first fine-grained benchmark specifically designed to assess action hallucinations in videos, comprising 11,306 video clips and 40,493 question-answer pairs across binary-choice, multiple-choice, and generative tasks. To mitigate affirmation bias, it introduces a bidirectional questioning protocol and bias-aware metrics. Experiments on ten state-of-the-art models reveal that hallucinations stemming from sequential reasoning are most severe, and that strong priors or fine-grained similarity significantly exacerbate the issue. Notably, high action recognition accuracy does not guarantee low hallucination rates, indicating a decoupling between recognition capability and hallucination robustness.
📝 Abstract
Video Large Language Models (VideoLLMs) have shown strong progress in video understanding, yet they still suffer from hallucinations that are inconsistent with visual evidence. Existing benchmarks mainly focus on object hallucination or coarse action perception, leaving a key video-specific problem underexplored: motion hallucination, in which models infer human motions that are absent from the video. We present MoHallBench, a benchmark for diagnosing motion hallucination in VideoLLMs. MoHallBench systematically evaluates three major sources of hallucination: co-occurrence priors, sequential inference, and similarity confusion. It contains 11,306 video clips and 40,493 question-answer pairs, covering binary-choice, multiple-choice, and generative settings. We further introduce a bi-directional questioning protocol with bias-aware metrics to reduce affirmation bias in binary evaluation. Experiments on ten recent open-source VideoLLMs reveal a clear decoupling between action recognition and hallucination resistance, as models that perform well on positive action recognition often fail on adversarial negatives. Among all settings, sequential inference hallucination is the most severe, showing that current models tend to over-infer expected outcomes from partial motion cues. Our analyses further confirm that stronger priors and finer-grained similarity substantially amplify hallucination. We hope MoHallBench can facilitate future evaluation and mitigation of motion hallucination in VideoLLMs.