CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection

📅 2026-05-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of existing video moment retrieval and highlight detection methods, which often overlook fine-grained intra-frame visual details guided by text, thereby constraining localization accuracy. To overcome this, the authors propose a unified spatiotemporal representation learning framework that, for the first time, jointly models text-driven progressive fine-grained image encoding and multi-scale temporal dynamics within an end-to-end architecture. By enhancing the collaborative alignment between frame-level spatial details and temporal context, the approach substantially improves video-text semantic matching. The method achieves state-of-the-art performance across four benchmark datasets: QVHighlights, Charades-STA, TACoS, and TVSum.
📝 Abstract
Video Moment Retrieval (MR) and Highlight Detection (HD) are crucial tasks in video analysis that aim to localize specific moments and estimate clip-wise relevance based on a given text query. Recent approaches treat them as similar video grounding tasks and use the same architecture to solve them. These tasks require both fine-grained comprehension at the image level and high-level temporal understanding across the entire video. Existing approaches have primarily focused on temporal modeling using frame-level features, often neglecting the rich visual information related to the text query within individual frames. This oversight leads to inaccurate grounding results. To address this limitation, we propose a Comprehensive Spatial-Temporal Representation Learning Framework (CoSTL), which captures both fine-grained image-level information and temporal dynamics. Specifically, CoSTL incorporates a text-driven progressive fine-grained image encoder, performing a two-step text-driven knowledge extraction process to learn fine-grained spatial representations. Furthermore, a multi-scale temporal perception module captures comprehensive spatial-temporal representations, enhancing the model's ability to process temporal dynamics. We demonstrate state-of-the-art performance on four public benchmarks: QVHighlights, Charades-STA, TACoS, and TVSum.
Problem

Research questions and friction points this paper is trying to address.

Video Moment Retrieval
Highlight Detection
Spatial-Temporal Representation
Fine-grained Visual Understanding
Text-Video Grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Comprehensive Spatial-Temporal Representation
Text-Driven Fine-Grained Encoding
Multi-Scale Temporal Perception
Moment Retrieval
Highlight Detection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xin Dong
Shenzhen International Graduate School, Tsinghua University
W
Wenjia Geng
Shenzhen International Graduate School, Tsinghua University
W
Wenfeng Deng
Pengcheng Laboratory
Y
Yansong Tang
Shenzhen International Graduate School, Tsinghua University