TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of identifying key factors in motion representation caused by the multivariate coupling of architectures, objectives, and data in video self-supervised learning. To this end, we propose the TT-VidT framework, which decouples spatial and temporal axes by initializing a Vision Transformer with DINOv3 and integrating a compact temporal transfer layer. This design acquires motion-centric representations by reconstructing target frames via differential compression. Extensive controlled experiments demonstrate that the synergy between TT3D and differential compression surpasses individual components, establishing a new paradigm for efficient, motion-sensitive pretraining. The proposed model achieves significant improvements across four action recognition benchmarks, including Jester, outperforming the strongest baselines by 54%–121% while substantially reducing encoder computational overhead.
📝 Abstract
Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized representations, whose gains concentrate on frame-to-frame change while retaining useful appearance. We address this with a matched $4 \times 6 = 24$ architecture-objective study at roughly 170M ~ 190M encoder scale on $\sim$1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, and propose TT-VidT. TT-VidT combines a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer, trained by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens. The sweep shows that TT3D with Diff Compression, not either component alone, enters the strongest motion-sensitive regime, and decoder ablations favor a compact video-pretrained decoder. In final comparison, TT-VidT leads Jester, Something-Something V2, ARID, and Diving48 fine-tuning simultaneously, improving over the strongest non-TT row by 54% ~ 121%, while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA2. HMDB51, IARD, and EPIC-Kitchens bound the claim.
Problem

Research questions and friction points this paper is trying to address.

Video Self-Supervised Learning
Motion-Centric Representation
Evaluation Methodology
Temporal Dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Pretraining
Motion-Centric Representation
Diff Compression
Temporal Decoupling
Self-Supervised Learning
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30