🤖 AI Summary
This study addresses why masked prediction captures features overlooked by unmasked reconstruction and how standard data augmentation obscures the advantages of dynamic masking. To investigate this, we construct a high-dimensional theoretical framework that decouples masking objectives from diversity benefits, demonstrating that masked linear reconstruction remains effective even when PCA fails, while quantifying the impact of mask diversity on sample complexity. These theoretical findings are validated through MAE-based modeling and experiments with CNN/ViT and BERT architectures. Our contributions reveal the statistical advantages of mask resampling, confirm that dynamic masking outperforms static alternatives and substantially enhances downstream task performance, and provide both theoretical and empirical guidance for optimizing pretraining pipelines.
📝 Abstract
Why can masked prediction learn useful representations that unmasked reconstruction misses? We study this question in a high-dimensional model of a masked autoencoder (MAE) trained on data with shared latent structure and heterogeneous noise. We prove that masked linear reconstruction can recover the latent feature at linear sample complexity in regimes where unmasked linear reconstruction, equivalent to PCA, fails. The analysis also quantifies the statistical advantage of mask resampling, an established ingredient of masked pretraining. By introducing a fixed collection of $K$ masks per sample, we characterize its effect on feature recovery and downstream performance, identifying regimes where greater mask diversity lowers sample complexity. Guided by this prediction, we find that random cropping and flipping in standard image-training pipelines can obscure the advantage of mask resampling by renewing the prediction task even when the patch mask is fixed. Removing these transformations reveals a downstream advantage for dynamic over static masking in CNN autoencoders and vision transformers. A complementary BERT pilot finds benefits from greater mask diversity on downstream language tasks. Our results separate the benefit of the masked prediction objective from that of mask diversity, and show how a tractable theory can guide experiments that uncover advantages hidden by standard training practices.