TSMD: Temporal-Stream Modality Dropout for Robust Video Highlight Detection

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited robustness caused by missing modalities in multimodal video highlight detection and the misalignment between mean squared error (MSE) loss and evaluation metrics. To tackle these issues, we propose a temporal-stream-level mixed modality dropout strategy that enhances model resilience to interference by simulating structured modality absence. Furthermore, we design a joint loss function integrating MSE, Pearson correlation, and RankNet to precisely align with the peak localization characteristics of highlights. Experimental results demonstrate that the proposed method achieves significant improvements in mAP@15 on the MoSu and Mr. HiSum datasets. Notably, it substantially outperforms baseline models under severe modality-missing conditions, validating its effectiveness and robustness for real-world multimodal video highlight detection scenarios.
📝 Abstract
Existing multimodal video highlight detectors typically assume that visual, audio, and textual streams are continuously available. In practice, however, inputs may suffer from localized frame missingness or complete-stream outage. We formulate this robustness challenge along two dimensions: temporal missingness, where frames are missing independently in each modality, and stream-level missingness, where one modality is unavailable throughout a video. Moreover, we find that the mean squared error (MSE) loss is misaligned with both the evaluation metrics and the peak-driven nature of highlights. Therefore, we propose Temporal-Stream Modality Dropout (TSMD), which combines structured missingness simulation with a joint objective comprising pointwise MSE, per-video Pearson correlation, and peak-oriented RankNet loss terms. TSMD has three variants: temporal, stream-level, and mixed dropout. On the MoSu and Mr. HiSum datasets, TSMD-Temporal improves mAP@15 by 7.06 and 3.41 points over TripleSumm under 50% independent temporal removal, whereas TSMD-Stream performs the best under complete-stream removal. TSMD-Mix retains most of these complementary benefits and ranks the best or the second-best across the evaluated temporal and stream-level conditions.
Problem

Research questions and friction points this paper is trying to address.

Video Highlight Detection
Multimodal Robustness
Temporal Missingness
Stream-level Missingness
Loss Misalignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Highlight Detection
Modality Dropout
Multimodal Robustness
Joint Objective Function
RankNet Loss
🔎 Similar Papers
2024-07-18IEEE Workshop/Winter Conference on Applications of Computer VisionCitations: 0
B
Bo-Yuan Cheng
AI Research Center, Inventec Corporation, Taiwan
K
Kuan-Yu Chen
AI Research Center, Inventec Corporation, Taiwan
P
Po-Han Huang
AI Research Center, Inventec Corporation, Taiwan
Jeng-Lin Li
Jeng-Lin Li
AI Center Lead of Inventec; Adjunct Assistant Professor of NTHU-BAI
machine learningreliable AImultimodalhealth analyticsspeech
J
Jian-Jiun Ding
Graduate Institute of Communication Engineering, National Taiwan University, Taiwan