Video-STLayout Pre-training

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过使用对比损失将视频特征与物体布局特征对齐,提出了一种新的预训练方法Video-STLayout,以提高复杂场景中的活动识别效果。
📝 Abstract
In recent years, pre-training has become fundamental to learning effective video representations, enabling strong transfer to downstream tasks. A popular framework in pre-training involves aligning features of a video encoder with that of another modality, for example, language or audio. We introduce Video-STLayout pre-training, a novel strategy for obtaining rich video representations informed by spatio-temporal layout of object bounding boxes. Object layouts can easily be obtained by applying an off-the-shelf object detector on the video frames. Our method uses a contrastive loss to align video features with the layout features from a trained layout encoder. We show the effectiveness of our approach in the task of activity recognition in complex scenes.
Problem

Research questions and friction points this paper is trying to address.

video representations
spatio-temporal layout
object bounding boxes
activity recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video-STLayout
spatio-temporal layout
object bounding boxes
contrastive loss
activity recognition
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30
💼 Related Jobs
No related jobs found.