🤖 AI Summary
This study addresses the limitation of existing DETR-based approaches in video moment retrieval and highlight detection, which fail to explicitly preserve query-relevant temporal evidence. To this end, we propose EviDETR, a framework that introduces an innovative evidence-preserving paradigm for efficient joint video understanding. Specifically, EviDETR incorporates semantic-aware feature reweighting, a sparse expert routing-based Mixture-of-Experts decoder, and a cross-task evidence fusion mechanism. The model leverages CLIP and SlowFast for dual-stream feature extraction, combined with confidence-weighted multiscale aggregation to enhance cross-task prediction consistency. Experimental results demonstrate that EviDETR achieves an R1@0.5 of 69.29 on the QVHighlights dataset and exhibits superior cross-dataset transferability across multiple benchmarks.
📝 Abstract
Joint video moment retrieval and highlight detection requires identifying query-relevant temporal segments while estimating clip-level saliency, yet DETR-style pipelines do not explicitly preserve query-relevant evidence throughout encoding, decoding, and cross-task prediction. We propose EviDETR, an evidence-preserving framework with three components. Semantic-aware Feature Reweighting (SFR) enhances query-relevant clip representations through saliency estimation and cross-modal interaction. A Temporal Top-2 Mixture-of-Experts (TTop2MoE) decoder performs query-adaptive refinement via sparse expert routing. MR-to-HD (MR2HD) fusion transfers span-level retrieval evidence to clip-level highlight prediction through confidence-weighted multi-scale aggregation. Using CLIP+SlowFast features, EviDETR achieves 69.29 R1@0.5, 54.77 R1@0.7, and 48.41 Avg. mAP for moment retrieval on QVHighlights, together with 41.83 HD-mAP and 68.33 HIT@1. Strong results on TACoS and Charades-STA further demonstrate cross-dataset transferability.