W2Rep: Learning Visual Representations by Watching the World Change

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent contradiction between image-based self-supervised learning, which lacks temporal information, and video-based methods that struggle to preserve single-frame feature independence. To resolve this, we propose a masked feature prediction framework that leverages video context and temporal difference conditioning. This approach assigns cross-frame targets a dual role: the image pathway learns time-invariant features, while the video pathway supplies missing evidence, enabling synergistic cross-frame prediction through independent and joint video encoding. Experiments demonstrate that the proposed framework significantly improves both frozen and fine-tuned recognition performance across various model scales. Furthermore, ablation studies validate the critical roles of directly updating source features and incorporating temporal shift mechanisms in achieving these gains.
📝 Abstract
Images capture the world at one moment, whereas video reveals how it changes. Image self-supervision learns spatial structure from a single moment, while video methods commonly learn temporal relationships inside a representation computed jointly from several frames. We ask whether watching a scene change can instead improve features available from one image without sacrificing the ability to represent video. We introduce W2Rep, a masked feature-prediction framework in which an independently encoded source image participates in prediction at the same or another moment. The predictor is conditioned on visible video context, the queried location, and the signed time interval between source and target. This gives the cross-frame objective two complementary roles: the image path learns features that remain useful across time, while the video path must gather evidence that is missing from the source image. Across model scales and downstream tasks, W2Rep improves frozen and fine-tuned recognition under our comparison protocol, while joint video encoding provides further gains over frame-wise aggregation. Controlled experiments show that these gains depend on directly updating the source-image features and on using both video context and temporal displacement. Overall, change across a video can supervise a visual encoder whose representations remain useful at either image or video granularity. Code is available at~\href{https://wenooi.github.io/W2Rep}{https://wenooi.github.io/W2Rep}.
Problem

Research questions and friction points this paper is trying to address.

visual representation learning
self-supervised learning
video understanding
image features
masked feature prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Masked Feature Prediction
Self-Supervised Learning
Video Representation
Cross-Frame Prediction
Visual Encoder
🔎 Similar Papers
No similar papers found.