🤖 AI Summary
This study addresses the limited robustness caused by missing modalities in multimodal video highlight detection and the misalignment between mean squared error (MSE) loss and evaluation metrics. To tackle these issues, we propose a temporal-stream-level mixed modality dropout strategy that enhances model resilience to interference by simulating structured modality absence. Furthermore, we design a joint loss function integrating MSE, Pearson correlation, and RankNet to precisely align with the peak localization characteristics of highlights. Experimental results demonstrate that the proposed method achieves significant improvements in mAP@15 on the MoSu and Mr. HiSum datasets. Notably, it substantially outperforms baseline models under severe modality-missing conditions, validating its effectiveness and robustness for real-world multimodal video highlight detection scenarios.
📝 Abstract
Existing multimodal video highlight detectors typically assume that visual, audio, and textual streams are continuously available. In practice, however, inputs may suffer from localized frame missingness or complete-stream outage. We formulate this robustness challenge along two dimensions: temporal missingness, where frames are missing independently in each modality, and stream-level missingness, where one modality is unavailable throughout a video. Moreover, we find that the mean squared error (MSE) loss is misaligned with both the evaluation metrics and the peak-driven nature of highlights. Therefore, we propose Temporal-Stream Modality Dropout (TSMD), which combines structured missingness simulation with a joint objective comprising pointwise MSE, per-video Pearson correlation, and peak-oriented RankNet loss terms. TSMD has three variants: temporal, stream-level, and mixed dropout. On the MoSu and Mr. HiSum datasets, TSMD-Temporal improves mAP@15 by 7.06 and 3.41 points over TripleSumm under 50% independent temporal removal, whereas TSMD-Stream performs the best under complete-stream removal. TSMD-Mix retains most of these complementary benefits and ranks the best or the second-best across the evaluated temporal and stream-level conditions.