HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection

๐Ÿ“… 2026-07-30
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge of modeling high-order cross-modal interactions among query text, contextual text, and local video frames in video misinformation detectionโ€”a task where existing methods often fail to preserve fine-grained veracity cues. To this end, we introduce, for the first time, a discriminative temporal hypergraph framework that represents multi-way dependencies among query tokens, evidence tokens, and sampled frames via a heterogeneous sparse hypergraph. The proposed architecture incorporates several key mechanisms: confidence-aware filtering, adaptive soft-association reasoning, residual text-video calibration, and discrepancy-aware readout, collectively preserving fine-grained spatiotemporal structure. Evaluated under the FactGuard protocol on FakeSV, FakeTT, and FakeVV datasets, our method achieves state-of-the-art accuracy of 83.7%, 82.0%, and 87.3%, respectively, significantly outperforming strong baselines while offering strong interpretability.
๐Ÿ“ Abstract
Video misinformation detection is often approached through global multimodal fusion or free-form multimodal reasoning. Both paradigms can under-represent localized authenticity cues that arise from coupled interactions among query phrases, contextual text, and short temporal spans of frames. Because such interactions are inherently higher-order, pairwise graph formulations are insufficient to capture multi-way cross-modal dependencies, whereas hypergraphs offer a suitable representation for these relations. We propose HyperClaim, a discriminative temporal hypergraph framework for sample-level authenticity classification. Using the title or benchmark-provided paired text as a claim-like query, HyperClaim constructs a sparse heterogeneous hypergraph over query tokens, evidence tokens, and sampled frames; applies confidence-aware filtering and source budgeting to form compact text-frame and short-range temporal evidence units; performs adaptive soft-incidence reasoning with residual text-video calibration; and aggregates textual, visual, and hyperedge states through a discrepancy-aware readout. Without relying on generated rationales or external tool calls, HyperClaim preserves fine-grained cross-modal and temporal structure that global fusion tends to flatten. Under the FactGuard temporal protocol, it achieves 83.7%, 82.0%, and 87.3% accuracy on FakeSV, FakeTT, and FakeVV, respectively, outperforming strong discriminative and reasoning-centric baselines. Learned incidence and attention weights further reveal token- and frame-level structure.
Problem

Research questions and friction points this paper is trying to address.

video misinformation detection
cross-modal reasoning
hypergraph
fine-grained authenticity cues
higher-order interactions
Innovation

Methods, ideas, or system contributions that make the work stand out.

hypergraph reasoning
cross-modal interaction
fine-grained misinformation detection
temporal evidence modeling
heterogeneous multimodal fusion
๐Ÿ”Ž Similar Papers
No similar papers found.