I Have a Stream: Making Self-Supervised Learning Work on Continuous Video

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation of self-supervised learning on continuous video streams caused by high intra-batch sample similarity, providing the first systematic analysis of streaming training bottlenecks. Methodologically, it constructs the WT++ dataset and proposes the StreamMAE framework, which optimizes the input pipeline via motion-biased cropping and introduces stream-aware regularization to improve the streaming adaptation of masked autoencoders (MAE). Experimental results demonstrate that the proposed approach surpasses existing streaming baselines, achieves performance comparable to independent and identically distributed (i.i.d.) training, and exhibits continuous improvement as data duration increases.
📝 Abstract
Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are independently sampled and globally shuffled across epochs. We study self-supervised learning from continuous video streams, where frames are consumed in temporal order using strict sliding-window batches, without global reshuffling or multi-epoch replay. To this end, we construct WT++, a 95-hour urban walking-tour video dataset for streaming pretraining. Combined with a comprehensive evaluation suite we find that contrastive and distillation-based methods struggle in this setting, while MAE is more robust but still falls short of standard i.i.d. pretraining. We find that high inter-batch similarity, caused by sliding-window consumption across consecutive batches, does not explain this gap. The main challenge is high intra-batch similarity, where frames within each batch are near-duplicates. To mitigate this, we propose StreamMAE, which preserves the core MAE reconstruction objective while adapting the input pipeline with stream-aware regularization and motion-biased crop selection. StreamMAE outperforms streaming baselines, matches i.i.d. MAE trained on the same video data, remains competitive with ImageNet-pretrained MAE, and scales positively as the pretraining stream grows from 12 to 95 hours.
Problem

Research questions and friction points this paper is trying to address.

Self-Supervised Learning
Continuous Video Streams
Streaming Pretraining
Intra-batch Similarity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Supervised Learning
Continuous Video Stream
Masked Autoencoder
StreamMAE
Sliding-Window Batches
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30