EviDETR: Preserving Query-Relevant Temporal Evidence for Moment Retrieval and Highlight Detection

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing DETR-based approaches in video moment retrieval and highlight detection, which fail to explicitly preserve query-relevant temporal evidence. To this end, we propose EviDETR, a framework that introduces an innovative evidence-preserving paradigm for efficient joint video understanding. Specifically, EviDETR incorporates semantic-aware feature reweighting, a sparse expert routing-based Mixture-of-Experts decoder, and a cross-task evidence fusion mechanism. The model leverages CLIP and SlowFast for dual-stream feature extraction, combined with confidence-weighted multiscale aggregation to enhance cross-task prediction consistency. Experimental results demonstrate that EviDETR achieves an R1@0.5 of 69.29 on the QVHighlights dataset and exhibits superior cross-dataset transferability across multiple benchmarks.
📝 Abstract
Joint video moment retrieval and highlight detection requires identifying query-relevant temporal segments while estimating clip-level saliency, yet DETR-style pipelines do not explicitly preserve query-relevant evidence throughout encoding, decoding, and cross-task prediction. We propose EviDETR, an evidence-preserving framework with three components. Semantic-aware Feature Reweighting (SFR) enhances query-relevant clip representations through saliency estimation and cross-modal interaction. A Temporal Top-2 Mixture-of-Experts (TTop2MoE) decoder performs query-adaptive refinement via sparse expert routing. MR-to-HD (MR2HD) fusion transfers span-level retrieval evidence to clip-level highlight prediction through confidence-weighted multi-scale aggregation. Using CLIP+SlowFast features, EviDETR achieves 69.29 R1@0.5, 54.77 R1@0.7, and 48.41 Avg. mAP for moment retrieval on QVHighlights, together with 41.83 HD-mAP and 68.33 HIT@1. Strong results on TACoS and Charades-STA further demonstrate cross-dataset transferability.
Problem

Research questions and friction points this paper is trying to address.

Moment Retrieval
Highlight Detection
DETR
Temporal Evidence
Video Understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Moment Retrieval
Highlight Detection
Mixture-of-Experts
Feature Reweighting
Evidence Fusion
🔎 Similar Papers
Haoran Sun
Haoran Sun
University of Electronic Science and Technology of China
LLMNLPAIML
Y
Yufan Li
Beijing Normal-Hong Kong Baptist University, Zhuhai, China
Q
Qichen Zhang
Beijing Normal-Hong Kong Baptist University, Zhuhai, China
H
Haoran Zhao
Beijing Normal-Hong Kong Baptist University, Zhuhai, China
S
Shuqi Wang
Beijing Normal-Hong Kong Baptist University, Zhuhai, China